ML Engineer, Inference Optimization at Build AI | CA, US | Rezi

ML Engineer, Inference Optimization at Build AI

ML Engineer, Inference Optimization

Build AI · CA, US

3 weeks ago

ML Engineer, Inference Optimization

Build AI · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this ML Engineer, Inference Optimization role.

Rezi rewrites your resume against Build AI's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the ML Engineer, Inference Optimization posting at Build AI — free, in seconds.

About the Role

We are hiring an Inference Optimization Engineer to make inference cheaper, faster, and good enough to scale our data engine and product without the GPU bill consuming the company.

Responsibilities

  • Own inference performance: latency, throughput, and cost per unit of work.
  • Optimize inference by focusing on kernels, batching, quantization, compilation, serving, and hardware utilization.
  • Profile pipelines using tools like Nsight or PyTorch Profiler to identify and resolve bottlenecks.
  • Collaborate with research and product teams to ensure models are both accurate and affordable to run at scale.
  • Build the serving and evaluation path to prevent experiments from obscuring inference costs.
  • Measure cost as a primary metric, integrated into the development process.

Requirements

  • Strong ML/systems engineering background with practical experience in inference optimization (serving, compilers, CUDA/kernels, quantization, or similar).
  • Proficiency in Python and C++ or Rust for performance-critical code.
  • Ability to think in terms of cost and performance metrics (dollars, tokens/frames per second) alongside accuracy.
  • Familiarity with PyTorch (or JAX) and profiling tools.
  • Comfort working in a small research team operating under cost constraints.

Skills

  • Inference optimization
  • Serving
  • Compilers
  • CUDA
  • Kernels
  • Quantization
  • Python
  • C++
  • Rust
  • PyTorch
  • JAX
  • Profiling tools
  • TVM
  • MLIR
  • TensorRT
  • GPU/accelerator cost management
  • Memory hierarchy understanding
  • Data movement optimization
  • Low-precision compute

Location

  • San Francisco (Financial District)
  • Shenzhen (Nanshan)

Work Type

  • In-person
  • Full-time

Experience Level

  • Mid-level
  • Senior

Benefits

  • Competitive pay
  • Medical, dental, and vision packages with generous premium coverage
  • $500 per month credit for waiving medical benefits
  • $2k per month housing subsidy for those living within walking distance of the office
  • Relocation support for those moving to San Francisco or Shenzhen
  • Various wellness benefits covering fitness, mental health, and more
  • Daily lunch and dinner in our office
  • Unlimited compute budget subject to ROI justification

About the Company

  • Build AI is the data hyperscaler for Physical AI, co-designing hardware, collection, infrastructure, and research to scale the physical labor dataset.
  • We believe in the Bitter Lesson and take a general approach of learning from humans, targeting all of physical labor.
  • We are a fully in-person team in San Francisco and Shenzhen, valuing engineering skills and expecting technical staff to contribute to both engineering and research across disciplines.

Equal Opportunity

  • Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply.