Architect and maintain high-performance storage infrastructure for AI training/inference workloads, building CSI drivers, GPUDirect Storage integrations, NVMe caching layers, and RDMA networking on Kubernetes.
You will architect the storage data-delivery fabric for intensive AI training and inference workloads. You will develop CSI drivers, integrate GPUDirect Storage, manage NVMe caching, tune I/O performance, optimize RDMA connectivity, implement storage monitoring and quotas, and enforce multi-tenant resource isolation.
Responsibilities
Design and maintain CSI drivers for high-performance parallel file systems
Implement GPUDirect Storage integrations
Develop local NVMe caching strategies for models and datasets
Optimize IOPS throughput and latency across the storage stack
Integrate storage with RDMA InfiniBand and RoCE networks
Implement storage performance monitoring and alerting
Define storage policies quotas and Kubernetes multi-tenancy isolation
Mentor engineers and lead architectural design reviews
Requirements
Bachelor’s or Master’s degree in Computer Science Electrical Engineering or a related field
5+ years of experience with distributed storage and high-performance file systems
Understanding of POSIX compliance and file I/O semantics
Deep expertise in Kubernetes CSI including volume plugins and storage operators
Experience with Linux block and file I/O and kernel performance tuning
Familiarity with RDMA InfiniBand and RoCE
Experience operating debugging and scaling production or HPC storage environments
Experience with Terraform Ansible and CI/CD pipelines