Founding Inference Engineer at General Compute | CA, US | Rezi

Founding Inference Engineer at General Compute

Founding Inference Engineer

General Compute · CA, US

1 weeks ago

Founding Inference Engineer

General Compute · CA, US

9 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Founding Inference Engineer role.

Rezi rewrites your resume against General Compute's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Founding Inference Engineer posting at General Compute — free, in seconds.

About the Role

Build and own the inference layer that sits between a bought-up model and a live customer request, handling request scheduling, batching, KV-cache management, and autoscaling across an ASIC fleet. This founding role offers wide scope and real ownership, requiring you to design the serving architecture and respond to incidents.

Responsibilities

  • Own the inference serving stack end-to-end, designing and building the system for production serving.
  • Push cost-per-token down by tuning batching strategy, KV-cache handling, and hardware utilization.
  • Build for reliability by implementing monitoring, alerting, and failover, and responding to incidents.
  • Work with the compiler and bring-up team to define interfaces and resolve issues.
  • Shape the roadmap by contributing to decisions on future serving features.
  • Set the technical bar for the growing serving team through early architecture and code-quality decisions.

Requirements

  • 5+ years building and operating production systems at the infrastructure layer, ideally including a high-throughput or low-latency serving system.
  • Direct experience with LLM inference serving (request batching, KV-cache management, continuous batching, or similar) in a production environment.
  • Comfortable owning reliability and designing for failure.
  • Strong systems fundamentals (concurrency, networking, scheduling) to reason about performance at the hardware level.
  • Self-directed and comfortable with ambiguity in a founding role.

Skills

  • LLM inference serving
  • Request batching
  • KV-cache management
  • Continuous batching
  • Production systems
  • Infrastructure layer
  • High-throughput serving
  • Low-latency serving
  • Reliability
  • Systems fundamentals
  • Concurrency
  • Networking
  • Scheduling
  • Hardware level performance
  • Serving frameworks (vLLM, TGI, TensorRT-LLM, SGLang)
  • Non-NVIDIA accelerators
  • Capacity planning
  • Fleet management

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Founding role
  • 5+ years

About the Company

  • General Compute is the neocloud for alternative chips.
  • We productionize purpose-built inference hardware from various manufacturers.
  • We manage the infrastructure (racks, data center space) for our customers.
  • Our hardware generates tokens 5–7× faster than GPU-based competitors.
  • Customers include frontier labs, AI application companies, and asset-light clouds.
  • Closed a $15M seed round in May 2026.
  • Closed a $400M debt facility with Upper90, collateralized by inference chips.