NPU Software Engineer Runtime
Rebellions · Seongnam-si, Gyeonggi-do, Korea ·
- Category
- Software engineering
- Experience
- 5+ years
You will build runtime modules that connect compilers, drivers, and ML frameworks for model deployment. You will support PyTorch execution, develop profiling capabilities, extend vLLM, optimize multi-NPU distributed inference, benchmark performance, and help deploy scalable inference services.
Responsibilities
- Design and implement runtime modules that interface with compilers and drivers
- Maintain native PyTorch execution support, torch.compile integration, and compiler toolchains
- Develop a user-facing performance profiler for the SDK
- Extend vLLM to improve NPU inference performance
- Design and optimize distributed multi-NPU inference and collective communication
- Benchmark, profile, and optimize runtime-system performance
- Deploy and scale inference services with ML and infrastructure engineers
Requirements
- Over 5 years of software-engineering experience with ML frameworks, inference runtimes, or AI accelerator toolchains
- Bachelor’s degree or higher in Computer Science, Electrical Engineering, or a related field
- Strong proficiency in C++ and Python
- Understanding of deep learning, LLM architectures, generative AI, and inference optimization
- Experience with LLM serving frameworks such as vLLM or TensorRT-LLM
- Understanding of tensor parallelism, KV-cache optimization, and memory-efficient execution
- Familiarity with compilers, runtimes, drivers, firmware, and hardware acceleration
- Debugging and performance-profiling skills for high-throughput inference
- Written and verbal communication skills