At mimic we are a frontier physical AI company pioneering general-purpose dexterous manipulation across the entire stack.
Spun out of ETH Zurich research in 2024, our team develops both state-of-the-art Video-Action Models and custom, in-house robotic hand hardware. Our growing team brings together world-class researchers working across our offices in Zurich and San Francisco.
The Role
As a Software Engineer – AI Infrastructure, you'll join a small, high-impact team building the compute, systems, and tooling that power our research and robots, from GPU clusters and training infrastructure to the platforms that deploy models onto real hardware. We're hiring engineers with different strengths, and you'll own meaningful parts of the stack from day one.
What you might work on
Depending on your background, you'll focus on one or more of these areas:
- Compute platform & DevOps: Build and operate our internal platform for managing GPU clusters and scheduling workloads. Own reliability, observability, security, access control, and cost efficiency across our compute.
- Distributed systems & scaling: Scale our infrastructure as data volume, training runs, and robot fleets grow. Design systems that stay fast and reliable under load, and help lay the foundations of our robot deployment infrastructure.
- Agentic systems: Build and scale LLM-based agentic workflows, including automated data labeling and orchestration of multi-step pipelines, with a focus on reliability, evaluation, and cost.
We don't expect you to cover all three areas. Deep expertise in one matters more than breadth across all of them.
Requirements
- Strong software engineering fundamentals: clean, modular, well-tested code, fluency with Git/GitHub, and a habit of proactively refactoring to keep systems understandable.
- Strong Python, plus comfort with at least one systems-level language or stack (e.g. Go, Rust, C++) or deep infrastructure tooling.
- Deep experience in at least one of: cloud/cluster infrastructure and DevOps, distributed systems, or production agentic/LLM systems.
- Clear communication: you explain technical decisions, collaborate across teams, and keep people informed without being asked.
- Ownership and initiative in a fast-moving, early-stage environment.
- Hands-on experience with cluster orchestration and scheduling, such as Kubernetes, Ray, Slurm, or similar.
- Genuine excitement about robots that learn from real-world data.
Nice to Have
- Terraform or other infrastructure-as-code.
- Managing GPU clusters (on-prem or cloud), including multi-node training setups.
- Security practices for infrastructure: networking, secrets management, IAM.
- Building agent frameworks, tool-use pipelines, or LLM evaluation systems.
- ML infrastructure fundamentals: PyTorch, data loading, distributed training, inference serving.
- Robotics, edge deployment, or sensor data (video, point clouds, time series).