Build and maintain a managed platform for AI model lifecycle, including fine-tuning, training pipelines, reinforcement learning, and experiment management for large language models.
You build a managed platform for the application development lifecycle with a focus on machine learning models and large language models. You manage fine-tuning systems, training pipelines, reinforcement learning workflows, and dataset, model, and experiment management at scale.
Responsibilities
Manage fine-tuning systems for large foundation models, including multi-node orchestration, checkpointing, failure recovery, and cost-efficient scaling
Implement and maintain end-to-end training pipelines for large language models
Apply reinforcement fine-tuning and reinforcement learning to fine-tuning and training workflows
Develop distillation and reinforcement learning pipelines for preference optimization, policy optimization, and reward modeling
Manage dataset, model, and experiment versioning, lineage, evaluation, and reproducible fine-tuning at scale
Requirements
Advanced degree in Computer Science, Engineering, or a related field
4-5+ years of industry experience leading impactful AI projects
Generative AI experience
Large language model experience
Multimodal model experience
Experience training, fine-tuning, and aligning large language models
Reinforcement learning experience
Reinforcement fine-tuning experience
Autonomous work
Collaboration skills
Proficiency in Golang or Python
PyTorch experience
Open-source AI project contributions
GPU performance optimization experience
Inference framework experience
Benefits
Restricted Stock Units
Paid time off
Paid holidays
Health insurance
Dental insurance
Vision insurance
Employer contributions to HSA account
Paid parental leave
Paid life insurance
Short-term disability insurance
Long-term disability insurance
Tuition reimbursement
Mental health and wellness support
Commuter benefits
Cell phone stipend
401(k) retirement plan with company match up to 4% of salary