Software Engineer - ML Infrastructure
Raydar · San Francisco, California, United States ·
- Category
- Software engineering
- Experience
- 6+ years
- Visa sponsorship
- Yes
- Salary
- USD 250,000 – 300,000 / year
Raydar · San Francisco, California, United States ·
Nebius Infra and Platform Early Talent Program · United States
MakerMaker · San Francisco
LinkedIn · Mountain View, CA, us
weekday-1 · India
ML infrastructure engineer at a healthcare AI imaging company, owning distributed training, reinforcement-learning, and production model-serving systems. Day to day: building training/checkpointing pipelines, RL rollout and reward-model infra, multimodal data pipelines, and serving/deployment tooling with Python, Kubernetes, Docker, and PyTorch or JAX.
About the company
Our client is a healthcare technology company developing AI-enabled imaging tools. Researchers and engineers work together on production machine learning systems for clinical applications.
The role
Raydar is recruiting for this role on behalf of our client. Own ML infrastructure across distributed training, reinforcement learning, and production serving. You will establish engineering practices and help researchers translate experiments into reliable systems.
What you'll do
- Build distributed training infrastructure with parallelism and checkpointing.
- Develop reinforcement-learning infrastructure for rollout generation, reward models, and experience collection.
- Partner with researchers to turn experimental workflows into production systems.
- Build data loading and preprocessing pipelines for multimodal datasets.
- Improve model serving, canary deployment, monitoring, and rollout tooling.
Requirements
What we're looking for
- 6+ years of ML infrastructure or distributed systems experience.
- Strong Python and hands-on production ML serving experience.
- Kubernetes and Docker experience with GPU scheduling, autoscaling, and reproducible environments.
- Distributed training expertise in PyTorch or JAX, including FSDP, DeepSpeed, or comparable approaches.
- Experience with RL or online-learning loops, logging, checkpointing, and evaluation.
- Breadth across the ML infrastructure stack plus technical leadership or mentoring experience.
Bonus points
- Startup experience.
- A/B testing and production ML experimentation platforms.
- A computer science or other STEM degree.
Benefits
Compensation and benefits
- Base salary: USD 250,000 to 300,000 per year.
- Equity: Competitive equity.
Location and work model
- San Francisco, California, United States.
- Five days per week onsite.
- Visa transfers and new visa sponsorships are supported.