DevOps / SRE Cloud Engineer (Site Reliability Engineer)
Recru-it · South Africa ·
- Seniority
- Senior
- Category
- Devops
- Experience
- 2+ years
Recru-it · South Africa ·
Recruit-It · South Africa
Hire Resolve · Cape Town, South Africa
WatersEdge Solutions · Cape Town, South Africa
Ciklum · Argentina
Senior DevOps/SRE Cloud Engineer to own an Azure-based Kubernetes platform. Responsibilities include managing production AKS clusters using Terraform, maintaining GitHub Actions CI/CD pipelines, and ensuring system reliability through Prometheus/Grafana observability and incident response.
SUMMARY:
Role purpose & context:
POSITION INFO:
Role purpose & context: Our client is looking for a senior DevOps \/ SRE Cloud Engineer to own the company's Azure-based Kubernetes platform end-to-end - resilient infrastructure, hardened CI\/CD, and strong observability so the company's workloads run reliably at scale. Key roles & responsibilities: · Operate production Kubernetes clusters and Azure infrastructure as Terraform code. · Maintain CI\/CD pipelines (GitHub Actions) with quality gates for reliable deployments. · Own observability: dashboards, alerting, SLOs, and incident response. · Manage secrets, identity, and network security across environments. · Partners with engineering teams to troubleshoot and improve reliability. Must-have technical skills \/ experience: · Azure cloud platform - production experience with AKS, ADLS Gen2 (blob storage), Azure Key Vault, Entra ID (incl. workload identity \/ managed identities), Azure networking (VNets, private endpoints, firewall rules), and Azure Service Bus or an equivalent message broker. · Kubernetes in production - deploying and operating self-managed (non-PaaS) workloads: Helm chart authoring, environment overlays, K8s operators, node-pool design, resource requests\/limits and capacity sizing, pod troubleshooting, upgrades. · Terraform - authoring and maintaining modular IaC for cluster, identity, storage, and secrets provisioning; state management and plan\/apply discipline across environments. · Docker - image builds (multi-arch), Compose-based local development stacks, container health checks and memory-limit tuning. · CI\/CD with GitHub Actions - building and maintaining pipelines for lint\/test\/build gates and container image publishing; branch-protection \/ required-status-check workflows. · Observability and SRE practice - Prometheus + Grafana (dashboards, alert rules), SLO\/alerting design, incident response, capacity planning from measured load; structured logging. · Linux and shell scripting - strong Bash; comfortable owning operational scripts as maintained, tested code. · Secrets management - provider-chain patterns (Key Vault in prod, env\/file locally); keeping secrets out of repos and images. Preferred \/ nice-to-have technical skills: · Apache Spark on Kubernetes operations - Spark Operator, cluster sizing, executor\/driver tuning, Spark Connect; Jupyter Hub on K8s. · Open data-platform stack - Hive Metastore, Trino, Apache Ranger, Delta Lake, S3-compatible object storage (MinIO); understanding of how governed query surfaces are wired. · Open Telemetry - SDK-based metrics\/traces instrumentation and collector topology; Open Lineage\/Marquez lineage. · Python - enough to maintain operational tooling and pytest-based environment\/infrastructure test tiers. · SQL Server and PostgreSQL - operational administration (backups, container deployments, EF-migration-driven schemas). · Multi-tenant \/ regulated-data environments - tenant isolation patterns, POPIA\/GDPR\/ISO 27001-adjacent controls; secret scanning and supply-chain hygiene (dependency audit, image provenance). Seniority and experience: · Senior (mid-senior acceptable with strong K8s depth). 5+ years DevOps\/SRE\/platform engineering, including 2-3 years running Kubernetes workloads in production and 2+ years on Azure. Must be able to own the environment independently. Required qualifications: · Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. · Cloud\/Kubernetes certification (e.g. Azure, CKA) advantageous. · Proven experience operating production-grade cloud infrastructure at scale; strong communication skills and ability to work independently in a hybrid\/ remote team. Why you'll love working for the company: The company believes in taking care of their team and creating an environment where you can thrive. As part of their company, you'll enjoy: · Flexible working arrangements: Whether you're a night owl or an early bird, they offer hybrid and remote options to suit your lifestyle · Comprehensive benefits: From a wellness program to home office reimbursements and continuous learning opportunities, they have got you covered. · Team culture: Fun team-building activities, regular socials, and a supportive, inclusive culture that values transparency, accountability, and work-life balance. · Performance incentives: Competitive salaries, ESOP, and recognition for your hard work.
Rakuten · RSIN_Indore