Model Bring-up Engineer / ML Compiler Engineer at General Compute | CA, US | Rezi

Model Bring-up Engineer / ML Compiler Engineer at General Compute

Model Bring-up Engineer / ML Compiler Engineer

General Compute · CA, US

2 weeks ago

Model Bring-up Engineer / ML Compiler Engineer

General Compute · CA, US

17 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Model Bring-up Engineer / ML Compiler Engineer role.

Rezi rewrites your resume against General Compute's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Model Bring-up Engineer / ML Compiler Engineer posting at General Compute — free, in seconds.

About the Role

You'll take a new model and get it running correctly on our ASIC in record time. You will own the loop from reference weights through the compiler to first correct tokens, building an agentic loop that compiles, runs, diffs, and localizes failures. This is a senior individual contributor role focused on owning the bring-up pipeline.

Responsibilities

  • Own model bringup end-to-end, from reference weights to first correct tokens running on ASIC within days.
  • Build the agentic bringup loop to automate compilation, execution, diffing, and failure localization.
  • Work within the compiler stack, addressing issues in graph capture, IR lowering, op coverage, and kernel selection.
  • Develop verification harnesses to ensure model correctness against reference implementations.
  • Optimize brought-up models for speed, including operator fusion, quantization, and memory layout.
  • Collaborate with the hardware partner's compiler and runtime team to deploy models in production.

Requirements

  • 5+ years in systems or ML systems, with expertise in ML compilers, model porting/bringup, or high-performance kernels.
  • Experience taking a self-designed model architecture and making it run correctly on a target it wasn't written for.
  • Proficiency in debugging numerical issues.
  • Strong understanding of modern LLM inference internals, including transformers, attention, KV cache, MoE routing, quantization, and batching.
  • Comfort working within compiler stacks like MLIR/LLVM, XLA, or vendor graph compilers at the IR, lowering, and op coverage levels.
  • Experience with agentic tooling for automating complex engineering loops.
  • Self-directed with the ability to anticipate and proactively address upcoming model bring-up tasks.

Skills

  • ML compilers
  • Model porting/bringup
  • High-performance kernels
  • Numerics debugging
  • LLM inference
  • Transformers
  • Attention
  • KV cache
  • MoE routing
  • Quantization
  • Batching
  • MLIR/LLVM
  • XLA
  • Vendor graph compilers
  • Agentic tooling
  • CUDA
  • Triton
  • Kernel development
  • Verification harnesses
  • Graph compilers
  • Serving runtimes

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Senior

About the Company

  • We are the first AI inference neocloud, using ASIC compute to generate tokens 5–7× faster than existing GPU-based competitors.
  • We recently closed an oversubscribed seed round and quadrupled our compute allocation to $97M.
  • Multiple lender conversations are live on a $200M asset-backed equipment facility.