Machine Learning Engineer - Infrastructure
Transparent Search Group · San Francisco, California, United States ·
- Employment
- Full time
- Category
- ML ai
- Experience
- 2+ years
- Visa sponsorship
- Yes
- Company size
- 11-50
- Salary
- USD 200,000 – 400,000 / year
Transparent Search Group · San Francisco, California, United States ·
4Minds · Dallas, TX
lilly · US, San Francisco CA
Eli Lilly and Company · US
Faculty AI · London
Infrastructure engineer at Causal Labs (AI startup) building and optimizing large distributed ML training and inference clusters for a Large Physics foundation Model, starting with weather prediction. Core stack: FSDP, DeepSpeed, NVIDIA GPUs, Python/C++, Kubernetes, Docker, GCP/AWS/Azure; on-site in San Francisco 5 days a week.
Company: Causal Labs
Location: San Francisco, CA (South Park office, in person 5 days per week; relocation provided)
Compensation: $200,000 - $400,000 + highly competitive early-stage equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers; can sponsor visas
Causal Labs is pursuing general causal intelligence: AI that can predict the future and identify the actions that change it. It is building a Large Physics foundation Model (LPM), because domains governed by physics have inherent cause-and-effect structure that visual or textual data lacks. Its starting domain is weather, the most observed physical system on earth, with rapid ground-truth feedback and data volumes that dwarf LLM training sets.
The founders come from Cruise, Google Research and Meta. The company is about 10 people in San Francisco, growing to around 35 this year, and is backed by Kindred Ventures, Refactor and BoxGroup.
Causal Labs is hiring infrastructure engineers to tackle the unsolved training and inference challenges of a Large Physics foundation Model. The work demands deep expertise in standing up distributed training clusters and optimizing performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology models and want to bet on a counterintuitive technical thesis, this is the role.
Initial call (30 min), technical screen, onsite day.
FSDP, DeepSpeed, NVIDIA GPUs, Python, C++, Linux, Kubernetes, Docker, GCP/AWS/Azure
Capital One · San Jose, CA