Member of Technical Staff, Developer Relations at Inferact | California, USA | Rezi

Member of Technical Staff, Developer Relations at Inferact

Member of Technical Staff, Developer Relations

Inferact · California, USA

1 months ago

Member of Technical Staff, Developer Relations

Inferact · California, USA

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Developer Relations Engineer to establish vLLM as the preferred method for developers to understand, build, and scale AI inference. This role requires an individual who can deeply comprehend vLLM as a complex LLM inference systems project, articulate challenging technical concepts clearly, and produce public resources that empower practitioners to develop superior systems.

Responsibilities

  • Write technical deep dives.
  • Build demos.
  • Create tutorials.
  • Contribute to documentation and examples.
  • Host workshops.
  • Educate developers on topics including KV cache, continuous batching, prefix caching, prefill and decode, quantization, GPU serving, latency versus throughput, and model-server tradeoffs across vLLM and adjacent systems.
  • Shape the AI infrastructure community's learning, adoption, and building practices with vLLM.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
  • Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.
  • Ability to clearly explain systems concepts such as KV cache, PagedAttention, continuous batching, prefill / decode scheduling, prefix caching, speculative decoding, tensor parallelism, data parallelism, or latency versus throughput tradeoffs.
  • Experience with vLLM or similar inference technologies (e.g., SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten-style serving platforms).
  • A strong public portfolio showcasing technical artifacts like blogs, tutorials, workshops, courses, OSS documentation, benchmark posts, architecture explainers, conference talks, demos, or runnable repositories.
  • Ability to write and teach for practitioners in a non-marketing style.
  • Strong engineering judgment, product taste, and the ability to transform raw technical material into valuable developer education.
  • Prior experience in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.
  • Experience creating technical content that teaches reusable mental models.
  • Experience contributing to developer-facing open source projects.
  • Existing credibility or community presence in AI infrastructure, OSS, CUDA / GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, Predibase, Together AI, Anyscale, LMSYS, or similar ecosystems.
  • Ability to host workshops, create hands-on labs, present technical talks, and guide developers from concept to working code.
  • Experience writing widely-shared technical blogs, courses, or architecture deep dives on LLM inference, model serving, GPU serving, or ML systems.
  • Experience building demos, benchmarks, tutorials, or repositories around vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, FlashInfer, or related systems.
  • Experience contributing to open-source ML infrastructure, inference systems, developer tooling, or technical education projects.
  • Experience creating practitioner-facing content with code, diagrams, benchmarks, demos, or end-to-end labs.
  • Experience building a durable personal portfolio demonstrating technical depth, taste, and a strong point of view.

Skills

  • LLM inference systems
  • Model serving
  • GPU inference
  • Distributed runtimes
  • Scheduling
  • Batching
  • Quantization
  • KV cache
  • PagedAttention
  • Continuous batching
  • Prefill / decode scheduling
  • Prefix caching
  • Speculative decoding
  • Tensor parallelism
  • Data parallelism
  • Latency versus throughput tradeoffs
  • vLLM
  • SGLang
  • TensorRT-LLM
  • TGI
  • LoRAX
  • Ray Serve
  • FlashInfer
  • BentoML
  • Baseten-style serving platforms
  • ML systems
  • Distributed systems
  • HPC
  • Compilers
  • GPU kernels
  • Serving infrastructure
  • MLOps
  • Developer tooling
  • Open-source infrastructure
  • CUDA / GPU
  • Ray
  • PyTorch
  • Modal
  • Predibase
  • Together AI
  • Anyscale
  • LMSYS

Location

  • San Francisco, California
  • Remote in US

Work Type

  • Remote
  • Onsite

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or similar.

Salary/Compensations

  • $200,000 - $400,000 USD + equity

Benefits

  • Generous health, dental, and vision benefits
  • 401(k) company match

About the Company

  • Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
  • Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.