Software Engineer GPU Inference
Cerebras Systems, Inc. · United States and Canada ·
- Category
- Software engineering
- Experience
- 5+ years
Cerebras Systems, Inc. · United States and Canada ·
NVIDIA · US, CA, Santa Clara
Cerebras · Toronto, CAN
Harell Data · Palo Alto, CA
NVIDIA · Canada, Toronto
Builds, deploys, and operates the GPU prefill path for Cerebras's AI inference service, working across API services, vLLM, PyTorch, ROCm, GPU nodes, and rack-scale infrastructure. Day to day, the engineer improves reliability, latency, throughput, and capacity through debugging, benchmarking, automation, and release validation.
You will build, deploy, and operate the GPU prefill path across API services, serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure. You will improve reliability, numerical correctness, observability, latency, throughput, and capacity efficiency through debugging, benchmarking, automation, and release validation.
NVIDIA · Canada, Toronto