Senior Site Reliability Engineer
Latitude · Remote ·
- Work mode
- Remote
- Seniority
- Senior
- Category
- SRE
Latitude · Remote ·
megaport · Sao Paulo
K2 Space · Los Angeles, CA
1global · Berlin, Berlin, Germany
Backblaze External Website · Remote - Bangalore
Build and maintain scalable, self-healing infrastructure using Kubernetes, Terraform, and observability tools like Prometheus and Grafana, while automating incident response and reliability processes.
You will build reliable, observable, and self-healing infrastructure at scale. You will automate operational tasks and incident response, improve monitoring and alerting, collaborate on resilient system designs, participate in on-call rotations, lead post-incident reviews, and document operational processes and runbooks.
gausslabs · Yeoksam, Seoul