Senior AI Data Center Network Engineer
Bitdeer · Singapore / Cyberjaya, Malaysia / Taipei, Taiwan ·
- Work mode
- Onsite
- Seniority
- Senior
- Category
- Network engineering
- Experience
- 10+ years
Bitdeer · Singapore / Cyberjaya, Malaysia / Taipei, Taiwan ·
NVIDIA · US, CA, Santa Clara
NVIDIA · US, CA, Santa Clara
NVIDIA · US, CA, Santa Clara
Lightning AI · New York, New York, United States; Remote; San Francisco, California, United States; Seattle, Washington, United States
You will architect high-availability network solutions for AI Cloud Data Centers, covering DCN, DCI, WAN, and backbone networks. You will monitor, troubleshoot, and tune performance for large-scale GPU clusters running on InfiniBand or RoCEv2 fabrics, and lead deep-dive investigations into complex issues affecting AI training and inference performance such as RDMA packet loss, latency, link flapping, and NCCL communication timeouts. You will manage NVIDIA Quantum and Spectrum series switches, ConnectX NICs, NetQ, and the UFM platform to keep the network fabric healthy and stable. You will lead network architecture changes, capacity expansions, cutovers, and firmware upgrades with zero incidents, and respond rapidly to critical incidents with immediate mitigation and thorough root cause analysis. You will build and maintain network monitoring and observability platforms, and develop automation tools to improve operational efficiency and standardize workflows.
NVIDIA · India, Bengaluru