Senior Software Engineer, AI Infrastructure
mirantis · Remote, USA, us ·
- Work mode
- Remote
- Seniority
- Senior
- Category
- Software engineering
- Experience
- 5+ years
mirantis · Remote, USA, us ·
JobCubby · Kentucky, United States
applied · Sunnyvale
Blockchain Association · Remote - Canada
ssv.network · Remote
Senior engineer on a small, remote-first team building Mirantis' new enterprise AI product for running and governing LLMs on customer Kubernetes clusters. Day to day: Go development of the model-serving layer — deployment, GPU scheduling and scaling, Helm-based enterprise packaging, integration with API/identity/metering services, and production observability for GPU inference.
Mirantis is building a new enterprise AI infrastructure product that lets organizations run and govern large language models on their own Kubernetes clusters. You will join a small senior team early, with broad ownership of the model-serving layer and its path to production.
What you'll do
Design and build LLM serving infrastructure on Kubernetes: deployment, GPU scheduling, scaling, and model lifecycle management.
Package the platform for enterprise environments: Helm-based installs, upgrades, and restricted/offline networks.
Integrate the serving layer with the platform's API gateway, identity, and metering services.
Build the observability for operating GPU inference in production (serving metrics, GPU telemetry).
Contribute across a multi-service codebase and help set engineering direction through design docs and reviews.
What we're looking for
5+ years of software engineering experience in infrastructure, platform, or distributed systems.
Deep hands-on Kubernetes experience: building and operating production workloads and Helm charts, not just consuming managed clusters.
Experience with GPU workloads or LLM inference, or strong adjacent systems experience and a track record of learning fast.
Strong Go programming skills; solid CI/CD and infrastructure-as-code skills.
Fluency with AI-assisted development tools (Claude Code, OpenAI Codex) as part of your daily engineering workflow.
Comfortable with high autonomy on a small, remote-first, written-culture team.
Nice to have
Inference performance work (quantization, batching, caching) or distributed serving frameworks.
Enterprise deployment experience: air-gapped installs, SSO/OIDC, supply-chain security.
UI development experience (e.g. React/TypeScript), useful as the product's management surfaces grow.
Open-source contributions in the Kubernetes or ML-infrastructure ecosystems.
What does Mirantis offer you?
We are a Leader for Container Management in G2 (#2 after AWS)!
NVIDIA · Israel, Yokneam