DevOps / SRE Cloud Engineer

Goldman Tech Resourcing (Pty) Ltd

Our client is seeking a senior DevOps / SRE Cloud Engineer to take ownership of an Azure-based Kubernetes environment, with a focus on resilient infrastructure, secure CI/CD, observability and reliable platform operations. The role is suited to an experienced engineer who can independently manage production environments and partner closely with development teams.

Key Responsibilities

  • Operate production Kubernetes clusters and Azure infrastructure using Terraform.
  • Maintain GitHub Actions CI/CD pipelines with appropriate quality gates.
  • Own observability through dashboards, alerting, SLOs and incident response.
  • Manage secrets, identity and network security across environments.
  • Partner with engineering teams to troubleshoot and improve system reliability.
  • Design and maintain modular Infrastructure as Code for clusters, identity, storage and secrets.
  • Support Kubernetes deployments, upgrades, capacity planning and operational troubleshooting.
  • Maintain secure and reliable containerised environments.
  • Develop and maintain operational Bash scripts as tested, maintainable code.
  • Apply effective secrets-management practices and prevent sensitive credentials from entering repositories or container images.

Requirements

  • 5+ years’ experience in DevOps, SRE or platform engineering.
  • 2–3 years’ experience operating Kubernetes workloads in production.
  • 2+ years’ production Azure experience.
  • Strong Azure experience with AKS, ADLS Gen2, Azure Key Vault, Entra ID, workload/managed identities, VNets, private endpoints, firewall rules and Azure Service Bus or an equivalent message broker.
  • Strong production Kubernetes experience, including Helm, environment overlays, Kubernetes operators, node-pool design, resource requests/limits, capacity sizing, troubleshooting and upgrades.
  • Strong Terraform experience, including modular IaC and state management across environments.
  • Strong Docker experience, including multi-architecture image builds, Compose, health checks and memory-limit tuning.
  • Hands-on GitHub Actions experience.
  • Strong Prometheus and Grafana experience, including dashboards and alert rules.
  • Understanding of SLOs, incident response and capacity planning.
  • Strong Bash and Linux skills.
  • Strong understanding of secrets management.
  • Bachelor’s degree in Computer Science, Engineering or equivalent practical experience.
  • Cloud or Kubernetes certification such as Azure certification or CKA would be advantageous.
  • Strong communication skills and the ability to work independently.

Advantageous Experience

  • Apache Spark on Kubernetes, Spark Operator and JupyterHub.
  • Hive Metastore, Trino, Apache Ranger, Delta Lake and MinIO.
  • OpenTelemetry and OpenLineage/Marquez.
  • Python and pytest-based infrastructure testing.
  • SQL Server and PostgreSQL administration.
  • Multi-tenant or regulated-data environments, including security and supply-chain controls.

Should you meet the requirements for this position, please email your CV to ***email_hidden***. You can also contact us on 031 350 4019 or alternatively you can visit our website https://stand-outstaffing.co.za

Correspondence will only be conducted with short listed candidates. Should you not hear from us within 4 days, please consider your application unsuccessful.