About the Role
We are seeking highly skilled and motivated software engineers to build AI inference systems that serve large-scale models with extreme efficiency. You will architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi-node, and multi-cloud environments. You will collaborate across inference, compiler, scheduling, and performance teams to push the frontier of accelerated computing for AI.
Responsibilities
- Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features.
- Profile and optimize the inference framework (vLLM) with methods like speculative decoding, data/tensor/expert/pipeline-parallelism, prefill-decode disaggregation.
- Develop, optimize, and benchmark GPU kernels (hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization.
- Build and extend high-level DSLs and compiler infrastructure to boost kernel developer productivity while approaching peak hardware utilization.
- Define and build inference benchmarking methodologies and tools.
- Contribute both new benchmark and NVIDIA’s submissions to the industry-leading MLPerf Inference benchmarking suite.
- Architect the scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds.
- Conduct and publish original research that pushes the pareto frontier for the field of ML Systems.
- Survey recent publications and find a way to integrate research ideas and prototypes into NVIDIA’s software products.
Requirements
- Bachelor’s degree (or equivalent experience) in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE) with 7+ years of experience; alternatively, Master’s degree in CS/CE/SE with 5+ years of experience; or PhD degree with the thesis and top-tier publications in ML Systems, GPU architecture, or high-performance computing.
- Strong programming skills in Python and C/C++.
- Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, distributed systems, deep learning theories.
- Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).
- Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL.
- Proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).
- Experience with containers and orchestration (Docker, Kubernetes, Slurm).
- Familiarity with Linux namespaces and cgroups.
- Excellent debugging, problem-solving, and communication skills.
- Ability to excel in a fast-paced, multi-functional setting.
Skills
- Python
- C/C++
- Go
- Rust
- CUDA
- NCCL
- Docker
- Kubernetes
- Slurm
- Linux namespaces
- cgroups
- ML frameworks
- inference engines
- GPU programming
- performance engineering
- profiling tools
- debug tools
- LLM inference engines
- ML compilers
- DSLs
- GPU libraries
- cloud platforms
- infrastructure as code
- CI/CD
- production observability
Location
- Hybrid
Work Type
- Hybrid
- Full-time
Experience Level
- 7+ years of experience with Bachelor's degree
- 5+ years of experience with Master's degree
- PhD with thesis and publications
Education Level
- Bachelor's degree in Computer Science, Computer Engineering, or Software Engineering
- Master's degree in Computer Science, Computer Engineering, or Software Engineering
- PhD degree
Salary/Compensations
- 170,000 CAD - 220,000 CAD for Level 4
- 225,000 CAD - 275,000 CAD for Level 5
Benefits
- Equity
About the Company
- At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work.
- Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential.
- Our team consists of experts in AI, systems and performance optimization.
- Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards.
- NVIDIA pioneered accelerated computing.
- Today, our AI infrastructure powers global intelligence, transforming every industry.
- Learn more about NVIDIA.
Equal Opportunity
- NVIDIA uses AI tools in its recruiting processes.
