Site Reliability Engineer (SRE) – GCP Platform
Kstack · São Paulo, SP, BR ·
- Work mode
- Hybrid
- Category
- SRE
- Experience
- 5+ years
Kstack · São Paulo, SP, BR ·
TENEX.AI · Remote, USA
cscgeneration-2 · Costa Rica
qadinc · Barcelona, CT, es
Synechron · Pune - Hinjewadi (Ascendas)
Build and maintain Google Cloud Platform infrastructure and Kubernetes clusters to automate deployments, improve reliability, and reduce manual toil for scalable, fault-tolerant systems.
About the job
Summary
Under general supervision, the Site Reliability Engineer (SRE) is responsible for improving the reliability, scalability, and resilience of digital platforms and infrastructure. Operating at the intersection of software engineering and systems engineering, this role focuses on building automation to eliminate toil (repetitive, manual tasks) and proactively preventing service-impacting incidents.
The SRE will design, build, and support large-scale, distributed, fault-tolerant systems in Google Cloud Platform (GCP). This role ensures that critical business platforms maintain high availability and performance while supporting a rapid pace of feature deployment. By championing observability and taking a holistic view of system health, the SRE will enhance cloud-based transformation initiatives, ensuring technology capabilities remain ahead of evolving customer needs and business growth.
Job Duties
Education & Experience
Knowledge, Skills, Abilities
Technical Skills (SRE & GCP Core):
Soft Skills & Process:
Physical Demands / Certifications
Devsu · Remote