Inference Engineer at Hyperbolic Labs | CA, US | Rezi

Inference Engineer at Hyperbolic Labs

Inference Engineer

Hyperbolic Labs · CA, US

1 weeks ago

Inference Engineer

Hyperbolic Labs · CA, US

14 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Inference Engineer role.

Rezi rewrites your resume against Hyperbolic Labs's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Inference Engineer posting at Hyperbolic Labs — free, in seconds.

About the Role

We are seeking an Inference Engineer to develop inference capabilities on Forge, our unified control plane. This role will enable customers to consume model tokens without managing GPUs and provide NeoCloud partners with a full-stack path to their own token-factory offering. You will be responsible for deploying and serving models across globally distributed clusters with heterogeneous hardware.

Responsibilities

  • Serve models on Forge and our Kubernetes offering.
  • Evaluate inference frameworks.
  • Establish monitoring, gateways, and endpoints for production readiness.
  • Optimize inference performance.
  • Implement autoscaling solutions.
  • Orchestrate KV-cache.
  • Debug customer inference issues.

Requirements

  • Strong general inference background with a broad understanding of the entire stack.
  • Deep Kubernetes experience, including production cluster operation.
  • Solid grasp of inference performance concepts like TTFT, disaggregated inference, speculative decoding, and KV cache.
  • Familiarity with modern inference frameworks and serving engines.
  • Working knowledge of NVIDIA Dynamo and its role in distributed serving architectures.
  • Experience setting up monitoring, gateways, and endpoints for production inference services.
  • Proven ability to build a product end to end, from inception to serving live traffic.
  • Strong self-initiative and comfort owning an area with minimal direction.
  • Generalist mindset, willing to take on adjacent tasks as needed by the product.

Skills

  • Inference deployment
  • Inference optimization
  • Model optimization (quantization, batching strategies, kernel-level tuning)
  • RDMA
  • High-performance networking
  • Distributed serving
  • Heterogeneous accelerator deployment
  • Customer support for inference debugging and performance issues
  • GPU cloud operations
  • Inference provider experience
  • AI infrastructure development

Location

  • Global

Work Type

  • Full-time

Experience Level

  • Primary seat for inference
  • Build it end to end

About the Company

  • Hyperbolic Labs is on a mission to democratize AI by making AI Cloud accessible through an Open-Access AI Cloud.
  • We utilize idle computing resources globally to offer an affordable and accessible GPU marketplace and AI inference service.
  • We are pioneers at the intersection of AI and open-source technology, advocating for a future where AI innovation is not limited by resource access.
  • We are looking for individuals passionate about making AI universally accessible, secure, and affordable.
  • Join us in building a platform that empowers innovators to realize their AI project visions.

Equal Opportunity

  • Hyperbolic is an equal opportunity employer.
  • We celebrate diversity and are committed to creating an inclusive environment for all employees.