ML Infrastructure Engineer - ML Compute Capacity
Apple · Santa Clara ·
- Category
- Devops
- Experience
- 7+ years
- Company size
- 1000+
Apple · Santa Clara ·
General Motors · Sunnyvale, California, United States of America
GRVTY · Springfield, Virginia, United States
Speechmatics · London, England, United Kingdom
JobCubby · Cork, Ireland
Engineer on Apple's ML Compute Capacity team builds and operates production systems that optimally distribute compute across Apple's large-scale accelerator fleet for ML training and inference. Work spans data pipelines, backend services (Python/Go), telemetry with Prometheus/Grafana, optimization algorithms, and tools on Kubernetes.
Scaling machine learning workloads across thousands of accelerators creates challenges that few engineers ever encounter. In Apple’s Machine Learning Platform Technologies organization, we build the infrastructure that powers large-scale ML training and inference workloads, bringing together expertise in distributed systems, machine learning infrastructure, and high-performance computing.
As an engineer on the ML Compute Capacity team, you will design, build, and operate the production systems that ensure compute resources are optimally distributed throughout the company. You'll work across the stack — from data pipelines and backend services to APIs and interactive frontends — developing telemetry systems, optimization algorithms, policies, and intuitive tools for managing demand and improving efficiency across Apple's largest accelerator fleet. Our small, nimble team works in a high-autonomy, fast-paced environment, and we're passionate about digging into data patterns, laying out the performance characteristics of an entire distributed system, and knowledge sharing. If the opportunity to own and operate services that scale, stay highly available, and "just work" excites you, then please reach out to us!
IMC Trading · Amsterdam, Netherlands; London, United Kingdom