You will build and operate software that turns large fleets of Cerebras systems, servers, and switches into reliable, observable clusters. You will automate bare-metal infrastructure, create Kubernetes operators and control-plane services, improve fleet reliability, and expose cluster capabilities through APIs, CLIs, and an MCP gateway.
Responsibilities
- Build declarative CRD-driven automation for bare-metal networking, operating systems, and application software across clusters
- Deliver push-button cluster installation, upgrades, and security patching with canaries and downtime budgets
- Develop Kubernetes operators for scheduling large inference workloads
- Build gRPC control-plane services, authorization, admission webhooks, and quota policies
- Create metrics and log pipelines, exporters, SLOs, and alerting for systems, servers, and network fabric
- Implement failure detection, highly available control planes, and automated recovery
- Develop CLIs, APIs, and an MCP gateway for users, operators, and AI agents
Requirements
- 5+ years building and operating production distributed systems or infrastructure software
- Production-quality Go and Python skills
- Experience writing or debugging Kubernetes controllers and operators
- Knowledge of CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC
- Debugging skills across distributed systems, Linux, and networking
- Experience with Prometheus, Grafana, PromQL, exporter design, and alerting
- Active use of coding agents and rigor in verifying their output