ML Infrastructure Engineer
Sunday Robotics · Redwood City, CA ·
- Work mode
- Onsite
- Category
- Devops
Builds and maintains the ML infrastructure behind robot learning: distributed model training on GPU clusters, multimodal robot data ingestion pipelines, and low-latency inference for real-time robot control, plus research tooling for debugging and experiment analysis. Key stack: SLURM/Kubernetes, distributed training, GPU performance optimization.
You will build systems across the robot learning pipeline, from multimodal data ingestion through distributed model training and real-time inference. You will maintain research infrastructure, optimize performance, manage datasets, and create tooling for debugging, visualization, and experiment analysis.
Responsibilities
- Maintain a research codebase for fast iteration and correctness
- Own model-training infrastructure, including scheduling, checkpointing, metrics, and logging
- Scale distributed training across GPU clusters
- Optimize memory usage and training throughput
- Build low-latency inference pipelines for real-time robot control
- Design pipelines for ingesting, validating, and transforming multimodal robot data
- Build storage and metadata indexing systems for datasets
- Build research tooling for debugging, visualization, and experiment analysis
Requirements
- Software engineering and systems fundamentals
- Experience building distributed systems or large-scale data pipelines
- Experience with ML training infrastructure
- Ability to reason about performance, memory, I/O, and GPU utilization
- Experience managing training workloads with SLURM, Kubernetes, or similar systems
- Ability to design, build, operate, and iterate on systems end to end
Benefits
- Equity