About the Role
vLLM is growing rapidly, and the CI system is critical to supporting this growth. This role focuses on ensuring the CI system scales effectively, maintaining and improving its performance and reliability.
Responsibilities
- Maintain and scale the compute infrastructure powering CI, release, performance benchmark, and accuracy evaluation for the vLLM project.
- Support a wide range of models and accelerators including H100/H200, (G)B200/300, AMD MI325/355X, TPU, Intel Gaudi, etc.
- Reduce CI time-to-signal from hours to minutes.
- Ensure comprehensive testing across the vLLM codebase.
- Maintain the stability and reliability of vLLM releases.
- Build tooling to support over 3,000 vLLM contributors.
Requirements
- Strong experience with Docker, Kubernetes, and containerized build or test environments.
- Experience building CI/CD pipelines from scratch using GitHub Actions, Buildkite, or similar systems.
- Familiarity with CI design patterns and techniques such as compute orchestration, handling flaky tests, dependency/environment management, caching, remote execution, test target determination, and test coverage.
- Proficiency in Python, Bash, Go, or similar languages for automation and tooling.
- Solid understanding of Linux, security, networking, storage, and package management.
- Experience setting up infrastructure for ML, inference, CUDA, ROCm, or accelerator-heavy workloads.
- Experience running Buildkite at scale, including agents, queues, dynamic pipelines, test sharding, caching, and artifact management.
- Experience operating Kubernetes clusters for CI, batch jobs, test execution, or internal developer infrastructure.
- Experience managing CI/CD in large open-source projects.
- Experience building dashboards, alerts, runbooks, or tooling for CI observability.
Skills
- Docker
- Kubernetes
- CI/CD
- GitHub Actions
- Buildkite
- Python
- Bash
- Go
- Linux
- ML Infrastructure
- Inference Infrastructure
- CUDA
- ROCm
- Accelerators
- Observability
Location
- San Francisco, California
- Remote (US)
Work Type
- Remote
- Onsite
Experience Level
- Mid-level
- Senior
Salary/Compensations
- $200,000 - $400,000 USD
Benefits
- Generous health, dental, and vision benefits
- 401(k) company match
About the Company
- Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
- Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
