Senior Software Engineer - AI Inference at NVIDIA | New York, USA | Rezi

Senior Software Engineer - AI Inference at NVIDIA

Senior Software Engineer - AI Inference

NVIDIA · New York, USA

2 weeks ago

Senior Software Engineer - AI Inference

NVIDIA · New York, USA

19 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Advance open-source LLM serving by contributing to inference engines like vLLM and SGLang, ensuring they run optimally on NVIDIA GPUs and systems. Improve the underlying stack for high-throughput, low-latency inference at scale. This is a hands-on role focused on performance bottlenecks, runtime improvements, and shipping community-beneficial changes.

Responsibilities

  • Contribute features, fixes, and optimizations upstream to vLLM/SGLang, including authoring PRs, participating in reviews, writing benchmarks/tests, and driving designs.
  • Implement and optimize inference-runtime capabilities such as batching, scheduling, streaming, request lifecycle management, and KV-cache efficiency.
  • Profile and improve performance-critical code paths across various layers, from Python to C++/CUDA, using data to guide optimization.
  • Enhance multi-GPU inference performance and reliability, focusing on parallelism, communication, and resource utilization.
  • Build and maintain performance and correctness regression tests to ensure stable behavior across different configurations.
  • Collaborate with model, platform, and SRE teams to translate production requirements into upstreamable solutions.

Requirements

  • 5+ years of experience building production software with strong systems engineering fundamentals.
  • Proven track record of delivering performance or reliability improvements.
  • Experience with LLM inference/serving stacks (e.g., vLLM, SGLang) and understanding performance tradeoffs.
  • Strong programming skills in Python and C++/CUDA, with the ability to debug and optimize performance-critical code.
  • Experience with profiling and performance investigation techniques (microbenchmarks, flame graphs, GPU profiling).
  • Familiarity with distributed systems concepts and concurrency.
  • Strong communication skills and experience working with open-source communities.
  • BS/MS in Computer Science, Computer Engineering, or related field, or equivalent experience.

Skills

  • Python
  • C++
  • CUDA
  • LLM inference/serving stacks
  • vLLM
  • SGLang
  • Profiling
  • Performance investigation
  • Distributed systems
  • Concurrency
  • Open-source collaboration

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Senior
  • 5+ years

Education Level

  • BS/MS in Computer Science, Computer Engineering, or related field, or equivalent experience.

Salary/Compensations

  • 152,000 USD - 241,500 USD for Level 3
  • 184,000 USD - 287,500 USD for Level 4

Benefits

  • Equity
  • Benefits

About the Company

  • NVIDIA pioneered accelerated computing and its AI infrastructure powers global intelligence, transforming every industry.
  • NVIDIA is considered one of the technology world's most desirable employers, with forward-thinking and creative people.
  • NVIDIA uses AI tools in its recruiting processes.

Equal Opportunity

  • NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal opportunity employer.
  • We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.