Head of Infrastructure at General Compute | CA, US | Rezi

Head of Infrastructure at General Compute

Head of Infrastructure

General Compute · CA, US

2 weeks ago

Head of Infrastructure

General Compute · CA, US

17 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Head of Infrastructure role.

Rezi rewrites your resume against General Compute's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Head of Infrastructure posting at General Compute — free, in seconds.

About the Role

You will own the infrastructure layer of our inference cloud end-to-end, managing the control plane, gateway, and observability stack. This role will evolve to encompass a heterogeneous fleet of ASICs and GPUs, requiring hands-on involvement with Kubernetes, dashboards, on-call duties, and direct collaboration with ASIC partners.

Responsibilities

  • Own the inference control plane, including modification, extension, and replacement of components.
  • Manage the gateway and load balancer for the fleet, focusing on model placement, request routing, and tail-latency engineering.
  • Own end-to-end observability, including per-request tracing, dashboards, SLOs, and alerting.
  • Perform capacity planning against a distributed traffic mix.
  • Manage the operational aspects of the ASIC partnership, serving as the technical point of contact for production issues.
  • Bring up the pre-fill side of the disaggregated architecture on a new hardware platform.
  • Build the on-call and incident response practice from scratch and grow the team.

Requirements

  • 7+ years in infrastructure, SRE, or platform engineering, with experience at an inference, ML, or HPC shop.
  • Hands-on experience with Kubernetes at production scale, including debugging complex issues.
  • Strong understanding of tail latency and its relation to utilization.
  • Comfortable managing vendor relationships where vendor issues become production problems.
  • Proven track record of building effective observability practices.
  • Experience with on-call duties and incident response.
  • Desire to be an early hire in a startup environment.

Skills

  • Kubernetes
  • Observability
  • On-call
  • Incident Response
  • Capacity Planning
  • Vendor Management
  • Infrastructure Management
  • SRE
  • Platform Engineering
  • MLOps
  • HPC
  • Tail Latency Optimization
  • Non-NVIDIA Accelerators (TPUs, ASICs, alternative GPUs)
  • Model Serving Stacks (vLLM, TGI, TensorRT-LLM, SGLang)
  • Network Fabric (RoCE, InfiniBand)
  • Team Hiring and Management
  • Hardware Boundary Understanding (firmware, drivers, thermals)

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • 7+ years

About the Company

  • We are the first AI inference neocloud, utilizing ASIC compute for faster token generation compared to GPU-based competitors.
  • We have secured significant funding, including an oversubscribed seed round and a $97M compute allocation.
  • We are in active discussions for a $200M asset-backed equipment financing facility.