Own a multi-GPU inference platform end to end: the operating system, storage, networking, observability, and model-serving stack. The role deploys and benchmarks open-weight models, manages routing, quotas, and GPU monitoring, and helps engineers integrate internal models, using Linux with Ansible, Terraform, Docker, and Kubernetes.
You will own a dedicated multi-GPU inference platform across its operating system, storage, networking, observability, model-serving stack, and user access. You will deploy and benchmark open-weight models, manage routing and quotas, document the platform, support integrations, and teach engineers how to use it effectively.
Responsibilities
Collaborate with IT to manage the GPU server lifecycle across operating systems, storage, and networking
Implement observability and alerting for GPU utilization, thermals, memory, latency, and throughput
Build and operate the model inference stack
Deploy, upgrade, and tune open-weight models
Evaluate and benchmark new models for quality, throughput, and latency
Recommend which models to run
Manage quotas, routing, and cost and usage reporting across teams
Optimize cloud model usage when needed
Support developers integrating internal models into IDE assistants, agents, CI pipelines, and applications
Maintain API keys, endpoints, and documentation
Write self-service onboarding guides
Run workshops and office hours
Design queueing and workload-prioritization approaches
Requirements
Have coding and Infrastructure as Code skills
Know Ansible, Terraform, Docker, or Kubernetes
Have Linux systems administration experience covering networking, storage, containers, systemd, and troubleshooting
Understand the open-weight model ecosystem
Communicate in English
Demonstrate a service-oriented approach to supporting engineers and documenting systems
Benefits
Option to be paid in bitcoin
Flexible working hours
Professional development budget for training, courses, and workshops