Leads customer LLM post-training engagements end to end, from data assessment through training and deployment. Daily work centers on building reinforcement-learning environments, synthetic-data and evaluation pipelines, running distributed training on GPU clusters, and mentoring engineers.
You will lead customer post-training engagements from data assessment through training and deployment. You will build reinforcement-learning environments, synthetic-data and evaluation pipelines, prepare datasets, run training loops, optimize distributed training, deploy models, and establish technical standards.
Responsibilities
Lead customer post-training engagements and define training strategies
Design reward signals, verifiers, datasets, training loops, and customer benchmarks
Build reinforcement-learning training environments
Build synthetic data generation, reward model, and preference data workflows
Build evaluation infrastructure and test sets
Own data pipelines from raw customer data to training-ready datasets
Deploy post-trained models across hybrid environments
Define playbooks and technical standards
Mentor engineers
Requirements
LLM post-training experience at scale
Reinforcement learning training environments
Preference optimization workflows
Synthetic data generation
Reward models
Verifiers
Evaluation methodology
Multi-node GPU clusters
Distributed training
Data pipelines
Version control for datasets and models
Reproducible workflows
6+ years of engineering experience
2+ years focused on LLM post-training in a leadership capacity
Customer-facing technical experience or interest
Benefits
Stock options
Medical, dental, vision, and life insurance
Annual wellness allowance
Daily office lunch and dinner
22 weeks of paid parental leave
Unlimited paid time off in the U.S.
30 vacation days in the U.K.
Visa sponsorship support
Regular off-sites, happy hours, and team celebrations