Builds and operates distributed training infrastructure for frontier-model pre-training across thousands of GPUs, optimizing throughput, stability, memory, and GPU utilization while maintaining dataset/checkpointing pipelines and debugging bottlenecks. Core tech includes Megatron, DeepSpeed, NCCL, and model parallelism.
You will build distributed systems for frontier-model pre-training and operate large-scale training runs. You will optimize throughput, stability, memory, communication, and GPU utilization; maintain pipelines for datasets and checkpointing; work with researchers to productionize experiments; and debug bottlenecks across training stacks and runtimes.
Responsibilities
Build and scale distributed training systems
Design and operate large-scale foundation model training runs
Develop infrastructure for training across thousands of GPUs
Optimize training throughput stability and efficiency
Productionize experimental training workflows with researchers
Improve communication memory usage and GPU utilization
Build and maintain training pipelines for datasets checkpointing and experiments
Debug distributed training and GPU performance bottlenecks