About the Role
Advance open-source LLM serving by contributing to inference engines like vLLM and SGLang, ensuring they run optimally on NVIDIA GPUs and systems. Improve the underlying stack for high-throughput, low-latency inference at scale. This is a hands-on role focused on performance bottlenecks, runtime improvements, and shipping community-beneficial changes.
Responsibilities
- Contribute features, fixes, and optimizations upstream to vLLM/SGLang, including authoring PRs, participating in reviews, writing benchmarks/tests, and driving designs.
- Implement and optimize inference-runtime capabilities such as batching, scheduling, streaming, request lifecycle management, and KV-cache efficiency.
- Profile and improve performance-critical code paths across various layers, from Python to C++/CUDA, using data to guide optimization.
- Enhance multi-GPU inference performance and reliability, focusing on parallelism, communication, and resource utilization.
- Build and maintain performance and correctness regression tests to ensure stable behavior across different configurations.
- Collaborate with model, platform, and SRE teams to translate production requirements into upstreamable solutions.
Requirements
- 5+ years of experience building production software with strong systems engineering fundamentals.
- Proven track record of delivering performance or reliability improvements.
- Experience with LLM inference/serving stacks (e.g., vLLM, SGLang) and understanding performance tradeoffs.
- Strong programming skills in Python and C++/CUDA, with the ability to debug and optimize performance-critical code.
- Experience with profiling and performance investigation techniques (microbenchmarks, flame graphs, GPU profiling).
- Familiarity with distributed systems concepts and concurrency.
- Strong communication skills and experience working with open-source communities.
- BS/MS in Computer Science, Computer Engineering, or related field, or equivalent experience.
Skills
- Python
- C++
- CUDA
- LLM inference/serving stacks
- vLLM
- SGLang
- Profiling
- Performance investigation
- Distributed systems
- Concurrency
- Open-source collaboration
Location
- Remote
Work Type
- Full-time
Experience Level
- Senior
- 5+ years
Education Level
- BS/MS in Computer Science, Computer Engineering, or related field, or equivalent experience.
Salary/Compensations
- 152,000 USD - 241,500 USD for Level 3
- 184,000 USD - 287,500 USD for Level 4
Benefits
- Equity
- Benefits
About the Company
- NVIDIA pioneered accelerated computing and its AI infrastructure powers global intelligence, transforming every industry.
- NVIDIA is considered one of the technology world's most desirable employers, with forward-thinking and creative people.
- NVIDIA uses AI tools in its recruiting processes.
Equal Opportunity
- NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal opportunity employer.
- We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
