Senior developer on NVIDIA's AI networking team building a highly optimized inference framework that runs on some of the world's largest supercomputers and data centers. Day-to-day work centers on modern C++/C/Rust, Linux internals, TCP/IP networking, and low-level performance tuning for next-generation AI infrastructure.
Development of a highly optimized inference framework
Running software on the world’s largest supercomputers and data centers
Working in a dynamic and challenging environment on innovative, next-generation products at the forefront of technology in terms of performance, scalability, and features
B.Sc. or equivalent experience in Computer Science or Software Engineering
6+ years of experience in modern C++ / C / Rust development
3 years of experience in Linux environment and familiarity with development tools
Deep knowledge of the TCP/IP network stack
Understanding of computer architecture and operating systems concepts
Background in Linux internals and low-level software optimizations including benchmarking, bottleneck research, and performance tuning
Experience in programming CUDA kernels
Familiarity with ML frameworks and LLMs
Background in parallel programming, high-performance computing, and RDMA technology