DevOps / SRE Cloud Engineer
Goldman Tech Resourcing (Pty) Ltd
Our client is seeking a senior DevOps / SRE Cloud Engineer to take ownership of an Azure-based Kubernetes environment, with a focus on resilient infrastructure, secure CI/CD, observability and reliable platform operations. The role is suited to an experienced engineer who can independently manage production environments and partner closely with development teams.
Key Responsibilities
- Operate production Kubernetes clusters and Azure infrastructure using Terraform.
- Maintain GitHub Actions CI/CD pipelines with appropriate quality gates.
- Own observability through dashboards, alerting, SLOs and incident response.
- Manage secrets, identity and network security across environments.
- Partner with engineering teams to troubleshoot and improve system reliability.
- Design and maintain modular Infrastructure as Code for clusters, identity, storage and secrets.
- Support Kubernetes deployments, upgrades, capacity planning and operational troubleshooting.
- Maintain secure and reliable containerised environments.
- Develop and maintain operational Bash scripts as tested, maintainable code.
- Apply effective secrets-management practices and prevent sensitive credentials from entering repositories or container images.
Requirements
- 5+ years’ experience in DevOps, SRE or platform engineering.
- 2–3 years’ experience operating Kubernetes workloads in production.
- 2+ years’ production Azure experience.
- Strong Azure experience with AKS, ADLS Gen2, Azure Key Vault, Entra ID, workload/managed identities, VNets, private endpoints, firewall rules and Azure Service Bus or an equivalent message broker.
- Strong production Kubernetes experience, including Helm, environment overlays, Kubernetes operators, node-pool design, resource requests/limits, capacity sizing, troubleshooting and upgrades.
- Strong Terraform experience, including modular IaC and state management across environments.
- Strong Docker experience, including multi-architecture image builds, Compose, health checks and memory-limit tuning.
- Hands-on GitHub Actions experience.
- Strong Prometheus and Grafana experience, including dashboards and alert rules.
- Understanding of SLOs, incident response and capacity planning.
- Strong Bash and Linux skills.
- Strong understanding of secrets management.
- Bachelor’s degree in Computer Science, Engineering or equivalent practical experience.
- Cloud or Kubernetes certification such as Azure certification or CKA would be advantageous.
- Strong communication skills and the ability to work independently.
Advantageous Experience
- Apache Spark on Kubernetes, Spark Operator and JupyterHub.
- Hive Metastore, Trino, Apache Ranger, Delta Lake and MinIO.
- OpenTelemetry and OpenLineage/Marquez.
- Python and pytest-based infrastructure testing.
- SQL Server and PostgreSQL administration.
- Multi-tenant or regulated-data environments, including security and supply-chain controls.
Should you meet the requirements for this position, please email your CV to ***email_hidden***. You can also contact us on 031 350 4019 or alternatively you can visit our website https://stand-outstaffing.co.za
Correspondence will only be conducted with short listed candidates. Should you not hear from us within 4 days, please consider your application unsuccessful.