Head of Engineering at Inferact | San Francisco, California, US | Rezi

Head of Engineering at Inferact

Head of Engineering

Inferact · San Francisco, California, US

Yesterday

Head of Engineering

Inferact · San Francisco, California, US

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Head of Engineering to build and lead the organization responsible for developing the systems that power vLLM and Inferact. This role requires an engineering leader with deep technical credibility in GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design to lead a senior-heavy, specialized engineering team.

Responsibilities

  • Build and lead the organization developing the systems that power vLLM and Inferact.
  • Partner closely with the founders to scale a senior-heavy, highly specialized engineering team.
  • Preserve the technical rigor, speed, and ownership that made vLLM successful.
  • Recruit and develop rare ML systems talent.
  • Translate ambitious research and infrastructure work into a focused execution plan.
  • Strengthen how teams operate.
  • Help Inferact deliver reliable, high-performance inference across models, hardware, and deployment environments.
  • Lead teams responsible for LLM serving, vLLM, SGLang, model execution, inference performance, GPU kernels, compiler or runtime systems, or distributed AI infrastructure.
  • Scale a small, senior-heavy engineering organization where the relevant talent market is narrow and technical quality matters more than headcount growth.
  • Integrate research-oriented or PhD talent into production teams, including setting expectations, structuring work, and building effective collaboration with product-focused engineers.
  • Represent the engineering organization credibly with open-source contributors, hardware partners, cloud providers, customers, candidates, and investors.
  • Build or lead engineering teams working directly on GPU or accelerator-level inference performance, ML compilers, kernels, runtimes, or hardware-software co-design.
  • Contribute to or lead teams around open-source ML systems projects such as vLLM, SGLang, PyTorch, Ray, Triton, XLA, ROCm, or related infrastructure.
  • Scale an engineering organization through an inflection point while preserving high technical standards, fast iteration, and direct ownership.
  • Recruit successfully from a global, highly competitive ML systems talent pool and build relationships with technical communities beyond traditional candidate pipelines.
  • Lead engineering in an early-stage AI infrastructure, developer infrastructure, distributed systems, or open-source company.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
  • Engineering leadership experience building and scaling highly specialized teams in LLM inference, ML systems, GPU or accelerator software, distributed systems, or closely related infrastructure.
  • Deep technical credibility at the inference layer, including hands-on understanding of inference runtimes, GPU or accelerator optimization, kernels, memory and communication bottlenecks, and hardware-software tradeoffs.
  • Ability to distinguish core inference-engine work from the routing, orchestration, and application layers above it, with opinions grounded in direct technical experience.
  • A strong record of recruiting, assessing, and retaining senior engineers, staff-level ICs, PhDs, and research-adjacent engineers in a production engineering environment.
  • Experience translating technically ambitious work into clear priorities, accountable ownership, execution plans, and durable engineering operating mechanisms.
  • Ability to remain close enough to the work to identify risks, pattern-match on difficult technical problems, and unblock teams without becoming a bottleneck or displacing technical ownership.
  • Strong judgment across organizational design, hiring, performance management, technical planning, execution cadence, and cross-functional decision-making.

Skills

  • vLLM
  • LLM inference
  • ML systems
  • GPU software
  • accelerator software
  • distributed systems
  • inference runtimes
  • GPU optimization
  • accelerator optimization
  • kernels
  • memory bottlenecks
  • communication bottlenecks
  • hardware-software tradeoffs
  • SGLang
  • model execution
  • inference performance
  • GPU kernels
  • compiler systems
  • runtime systems
  • distributed AI infrastructure
  • ML compilers
  • PyTorch
  • Ray
  • Triton
  • XLA
  • ROCm

Location

  • San Francisco, California

Work Type

  • Relocation considered for exceptional candidates

Experience Level

  • Engineering leadership experience
  • Experience scaling a small, senior-heavy engineering organization
  • Experience integrating research-oriented or PhD talent into production teams
  • Experience leading teams responsible for LLM serving
  • Experience leading engineering in an early-stage AI infrastructure, developer infrastructure, distributed systems, or open-source company.

Education Level

  • Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.

Salary/Compensations

  • Highly competitive base and meaningful equity

Benefits

  • Generous health, dental, and vision benefits
  • 401(k) company match

About the Company

  • Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
  • Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

Equal Opportunity

  • We sponsor visas on a case-by-case basis.