Own server-side platform architecture for Cerebras AI clusters: define server configurations and baselines (CPU, memory, PCIe, NIC, NVMe, BIOS/firmware), build capacity and performance models, benchmark and qualify vendor platforms, and support deployment debugging. Core stack is x86 servers, Linux, RDMA/RoCE networking, and C/C++/Python.
You will own server-side platform architecture for AI clusters. You will define server configurations, capacity formulas, CPU, memory, PCIe, networking, storage, and firmware baselines. You will model and benchmark performance, lead vendor engagements, establish qualification criteria, and troubleshoot deployment regressions.
Responsibilities
Own architecture for cluster server roles, configurations, and lifecycle strategy
Define server formulas, capacity planning, and headroom policy
Specify CPU, memory, PCIe, NIC, and NVMe platform configurations
Translate runtime flows into hardware requirements
Develop and validate performance and scaling models
Define operating system, BIOS, firmware, and driver baselines
Evaluate emerging server technologies
Lead vendor engagements and qualification efforts
Support deployment debugging and root-cause analysis
Requirements
PhD and 8+ years of industry experience, or BS or MS and 10+ years of industry experience
5+ years of server platform architecture, systems performance engineering, or large-scale infrastructure design