Software engineer who builds and operates the platform managing Cerebras' large fleets of AI clusters — services, integrations, and tools that give operators visibility into cluster health, capacity, performance, and incidents. Core stack: Go or Python, Kubernetes, Linux, containers, and distributed-systems engineering.
You will build and operate software that manages large fleets of AI clusters. You will create services, integrations, operational tools, and user-facing applications that help operators monitor health, capacity, performance, and incidents. You will lead projects from design through production, automate operational workflows, and improve reliability as the fleet grows.
Responsibilities
Build and operate software for managing large fleets of AI clusters
Provide operators with actionable views of cluster health, capacity, performance, and issues
Develop services and integrations across infrastructure systems
Automate incident investigation and service-restoration workflows
Design reliable systems that withstand component and site failures
Gather platform-user needs and make practical product and engineering decisions
Lead projects from design through production and use operational feedback to improve them
Requirements
12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure
Strong Go or Python skills
Experience designing services and APIs
Expertise in control planes, fleet management systems, or operational platforms
Experience with Linux, containers, Kubernetes, and distributed-system failures
Experience designing for asynchronous work, retries, and partial failures
Experience with event streaming, workflow automation, or time-series telemetry
Strong judgment in reliability, security, and observability
Ability to lead ambiguous projects and collaborate across engineering and operations teams