Technical program manager role at Cerebras overseeing capacity planning, forecasting, allocation, and utilization for AI inference clusters. Day to day this means coordinating customer deployments, model launches, and datacenter bring-up, driving incident resolution and postmortems, and using tools like SQL, Grafana, Python, Jira, and Confluence.
You will lead capacity planning, forecasting, allocation, and utilization reporting for inference clusters. You will coordinate customer deployments, model launches, datacenter bring-up, and capacity-management tool adoption; maintain execution records; and drive risk mitigation, incident resolution, and postmortems.
Responsibilities
Run capacity planning, deployment tracking, utilization reporting, and forecasting
Plan capacity for customer deployments and model launches
Support new datacenter capacity bring-up and production readiness
Coordinate cluster allocation, model placement, and rebalancing
Drive adoption and improvement of capacity-management tools
Identify and mitigate capacity risks and bottlenecks
Lead capacity-related incident resolution and postmortems
Maintain Jira epics and Confluence documentation
Requirements
Five or more years of technical program management or product operations experience in cloud infrastructure, ML serving, or capacity planning
Experience leading cross-functional engineering, product, and operations programs
Knowledge of inference serving, model replicas, batching, prefill, decode, KV cache, and accelerator scheduling