Principal Software Engineer, Distributed Systems
Sigma Automate · US ·
- Seniority
- Principal
- Category
- Software engineering
Sigma Automate is hiring a Staff/Principal-level Site Reliability Engineer focused on distributed systems and platform reliability. The role owns the architecture and engineering practices that keep Sigma workflows running reliably despite component failures such as crashes, timeouts, duplicate messages, and Kubernetes pod restarts.
Background We are hiring a Staff / Principal Site Reliability Engineer focused on distributed systems and platform reliability . This is not a traditional DevOps role. You will own the architecture and engineering practices that ensure Sigma workflows execute reliably even when individual components fail. Processes crash. Networks become unavailable. APIs time out. Workers hang. Messages may be delivered more than once. Databases experience contention. Kubernetes pods restart. Your job is to de…