Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Founding Inference Engineer role.
Rezi rewrites your resume against General Compute's job description. Free.

Tailor your resume to this Founding Inference Engineer role.
Rezi rewrites your resume against General Compute's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Founding Inference Engineer posting at General Compute — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Founding Inference Engineer posting at General Compute — free, in seconds.
About the Role
Build and own the inference layer that sits between a bought-up model and a live customer request, handling request scheduling, batching, KV-cache management, and autoscaling across an ASIC fleet. This founding role offers wide scope and real ownership, requiring you to design the serving architecture and respond to incidents.
Responsibilities
- Own the inference serving stack end-to-end, designing and building the system for production serving.
- Push cost-per-token down by tuning batching strategy, KV-cache handling, and hardware utilization.
- Build for reliability by implementing monitoring, alerting, and failover, and responding to incidents.
- Work with the compiler and bring-up team to define interfaces and resolve issues.
- Shape the roadmap by contributing to decisions on future serving features.
- Set the technical bar for the growing serving team through early architecture and code-quality decisions.
Requirements
- 5+ years building and operating production systems at the infrastructure layer, ideally including a high-throughput or low-latency serving system.
- Direct experience with LLM inference serving (request batching, KV-cache management, continuous batching, or similar) in a production environment.
- Comfortable owning reliability and designing for failure.
- Strong systems fundamentals (concurrency, networking, scheduling) to reason about performance at the hardware level.
- Self-directed and comfortable with ambiguity in a founding role.
Skills
- LLM inference serving
- Request batching
- KV-cache management
- Continuous batching
- Production systems
- Infrastructure layer
- High-throughput serving
- Low-latency serving
- Reliability
- Systems fundamentals
- Concurrency
- Networking
- Scheduling
- Hardware level performance
- Serving frameworks (vLLM, TGI, TensorRT-LLM, SGLang)
- Non-NVIDIA accelerators
- Capacity planning
- Fleet management
Location
- Remote
Work Type
- Full-time
Experience Level
- Founding role
- 5+ years
About the Company
- General Compute is the neocloud for alternative chips.
- We productionize purpose-built inference hardware from various manufacturers.
- We manage the infrastructure (racks, data center space) for our customers.
- Our hardware generates tokens 5–7× faster than GPU-based competitors.
- Customers include frontier labs, AI application companies, and asset-light clouds.
- Closed a $15M seed round in May 2026.
- Closed a $400M debt facility with Upper90, collateralized by inference chips.