Member of Technical Staff - ML Infra
Trajectory · San Francisco ·
- Work mode
- Onsite
- Employment
- Full time
- Category
- Software engineering
Trajectory · San Francisco ·
Blackrock Neurotech · Salt Lake City, UT
featherlessai · Remote (world)
Reducto · San Francisco Office
Build and optimize ML infrastructure powering self-improving AI systems across distributed training/RL, inference serving, and GPU kernels. Day-to-day work involves GPU scheduling, checkpointing, KV-cache and batching optimization, and CUDA/Triton kernel development alongside researchers.
As an ML Infrastructure Engineer at Trajectory, you will develop algorithms and own the innermost loop of our continual learning platform: the ML training stack!
Our team’s goal is to improve the training stack by 10× every month across GPU scale, model size, throughput, memory efficiency, caching, and latency.
This role spans training, inference, and kernels. Bring deep expertise in at least one area and curiosity across the stack. We’ll shape your initial ownership around your strengths and have high opportunities for you to work across the whole ML training platform.
Training
Build and optimize distributed training and RL infrastructure, including rollout execution, GPU scheduling, training parallelisms, checkpointing, and recovery. Improve training experimentation throughput and shorten research iteration cycles is a key part of training 100B+ parameter models here.
Inference
Optimize serving for production agents and training rollouts. Improve batching, scheduling, and KV-cache management while balancing latency, throughput, cost, and model quality.
Kernels and runtimes
Develop and optimize GPU kernels and runtimes using CUDA, Triton, or comparable tools. Improve memory use and execution efficiency, preserving numerical correctness and verifying gains in real workloads.
Across all three
Build reproducible benchmarks, observability, and automated research workflows that propose changes, run experiments, and validate improvements. Work with researchers to turn new algorithms into reliable systems across training, inference, and kernels.
We’re building one of the world’s best continual learning loops for ML infrastructure: agents that can propose optimizations, run experiments, measure gains, and learn from experiments across training, inference, and kernels. Come join us to accelerate continual learning’s inner loop!
Strong fundamentals in distributed systems, networking, storage, and failure recovery, with experience shipping demanding systems.
Required deep specialization in one of training infrastructure, inference systems, or GPU kernels, supported by systems built or measurable optimizations delivered.
Strong Python skills and languages relevant to your specialty, such as C++, CUDA, or Triton.
Understanding of PyTorch, JAX, or comparable framework internals, with strong profiling and debugging skills.
Ownership and clear communication: work closely with researchers and deliver measurable performance gains while preserving correctness and reliability.
We value demonstrated capability over credentials.
Hands-on experience with training and RL stacks such as Miles, SkyRL, Prime Intellect’s verifiers, or comparable systems. Depending on your specialty, experience with vLLM, SGLang, collective communication, or ML compilers is also valuable.
Trajectory is a research and product lab creating the platform for continual learning.
AI is the most capable software ever built. Every valuable correction and edit that happens in a product evaporates at the next session. A few teams have closed this gap by hand-coupling their models to their products: Composer, Claude Code, Windsurf SWE-1.
Trajectory is the first scalable approach for every company: our platform unlocks the signal already sitting in product use, so companies can continuously post-train large-scale agentic models that outperform the frontier.
Our research team comes from Deepmind, OpenAI, Meta Superintelligence, and product team from Figma, Apple, Stripe, and Windsurf. We're working with customers like Harvey, Rogo, Mercor, Decagon and Clay, and we've raised $60M from Sequoia, Conviction, Jeff Dean, and Fei Fei Li.
Inception · San Mateo, USA