About the Role
We are seeking a Developer Relations Engineer to establish vLLM as the preferred method for developers to understand, build, and scale AI inference. This role requires an individual who can deeply comprehend vLLM as a complex LLM inference systems project, articulate challenging technical concepts clearly, and produce public resources that empower practitioners to develop superior systems.
Responsibilities
- Write technical deep dives.
- Build demos.
- Create tutorials.
- Contribute to documentation and examples.
- Host workshops.
- Educate developers on topics including KV cache, continuous batching, prefix caching, prefill and decode, quantization, GPU serving, latency versus throughput, and model-server tradeoffs across vLLM and adjacent systems.
- Shape the AI infrastructure community's learning, adoption, and building practices with vLLM.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
- Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.
- Ability to clearly explain systems concepts such as KV cache, PagedAttention, continuous batching, prefill / decode scheduling, prefix caching, speculative decoding, tensor parallelism, data parallelism, or latency versus throughput tradeoffs.
- Experience with vLLM or similar inference technologies (e.g., SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten-style serving platforms).
- A strong public portfolio showcasing technical artifacts like blogs, tutorials, workshops, courses, OSS documentation, benchmark posts, architecture explainers, conference talks, demos, or runnable repositories.
- Ability to write and teach for practitioners in a non-marketing style.
- Strong engineering judgment, product taste, and the ability to transform raw technical material into valuable developer education.
- Prior experience in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.
- Experience creating technical content that teaches reusable mental models.
- Experience contributing to developer-facing open source projects.
- Existing credibility or community presence in AI infrastructure, OSS, CUDA / GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, Predibase, Together AI, Anyscale, LMSYS, or similar ecosystems.
- Ability to host workshops, create hands-on labs, present technical talks, and guide developers from concept to working code.
- Experience writing widely-shared technical blogs, courses, or architecture deep dives on LLM inference, model serving, GPU serving, or ML systems.
- Experience building demos, benchmarks, tutorials, or repositories around vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, FlashInfer, or related systems.
- Experience contributing to open-source ML infrastructure, inference systems, developer tooling, or technical education projects.
- Experience creating practitioner-facing content with code, diagrams, benchmarks, demos, or end-to-end labs.
- Experience building a durable personal portfolio demonstrating technical depth, taste, and a strong point of view.
Skills
- LLM inference systems
- Model serving
- GPU inference
- Distributed runtimes
- Scheduling
- Batching
- Quantization
- KV cache
- PagedAttention
- Continuous batching
- Prefill / decode scheduling
- Prefix caching
- Speculative decoding
- Tensor parallelism
- Data parallelism
- Latency versus throughput tradeoffs
- vLLM
- SGLang
- TensorRT-LLM
- TGI
- LoRAX
- Ray Serve
- FlashInfer
- BentoML
- Baseten-style serving platforms
- ML systems
- Distributed systems
- HPC
- Compilers
- GPU kernels
- Serving infrastructure
- MLOps
- Developer tooling
- Open-source infrastructure
- CUDA / GPU
- Ray
- PyTorch
- Modal
- Predibase
- Together AI
- Anyscale
- LMSYS
Location
- San Francisco, California
- Remote in US
Work Type
- Remote
- Onsite
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or similar.
Salary/Compensations
- $200,000 - $400,000 USD + equity
Benefits
- Generous health, dental, and vision benefits
- 401(k) company match
About the Company
- Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
- Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
