A staff-level SRE who designs and operates reliable, scalable software delivery and platform infrastructure across datacenters and clouds — building self-service workflows, defining SLOs and error budgets, automating operational toil, supporting incidents, and mentoring other SREs. Core stack includes GitOps/Argo CD CI/CD and Prometheus, Loki, Tempo, and Mimir observability.
You will lead the design of reliable, scalable software delivery and operational platforms. You will build self-service workflows, evolve reliability practices, automate operational toil, support incident escalations, mentor SREs, and measure improvements in deployment velocity, reliability, and service ownership.
Responsibilities
Define and implement a strategy for reliably delivering and operating software across multiple datacenters and cloud solutions
Architect self-service platforms and internal tooling for critical workflows
Define and evolve SLOs, SLIs, error budgets, postmortems, chaos testing, and capacity forecasting
Mentor SREs and prioritize automation based on production pain points