Performance Engineer (Inference, Training & GPU) at World Labs | CA, US | Rezi

Performance Engineer (Inference, Training & GPU) at World Labs

Performance Engineer (Inference, Training & GPU)

World Labs · CA, US

3 weeks ago

Performance Engineer (Inference, Training & GPU)

World Labs · CA, US

23 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Performance Engineer (Inference, Training & GPU) role.

Rezi rewrites your resume against World Labs's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Performance Engineer (Inference, Training & GPU) posting at World Labs — free, in seconds.

About the Role

World Labs is seeking a Performance Engineer to optimize the speed of their AI models, ensuring they run as fast as the hardware allows. This role involves identifying and eliminating performance bottlenecks across the entire system, from low-level kernel optimization to fleet-wide serving efficiency, in close collaboration with researchers.

Responsibilities

  • Optimize inference and serving end to end — latency, throughput, batching, caching, and scheduling — to serve our models efficiently at production scale.
  • Write and tune GPU kernels (CUDA, Triton) for hot paths; drive kernel fusion, memory- and bandwidth-bound optimization, and low-precision (FP8/INT8) execution.
  • Optimize training throughput and GPU utilization: parallelism strategies, communication/compute overlap, mixed precision, and eliminating pipeline stalls.
  • Build performance models, profiling workflows, and observability that make throughput, latency, cost, utilization, and their tradeoffs legible across the stack.
  • Own numerical correctness across precision, kernel, and hardware changes — treating correctness as part of performance, not separate from it.
  • Partner with researchers to productionize models for serving and to make experiments run faster and more reliably.
  • Work on the distributed systems that training and inference run on, focusing on maximizing GPU efficiency.

Requirements

  • Strong performance-engineering foundations: profiling, roofline analysis, latency/throughput optimization, and disciplined root-cause investigation.
  • Deep GPU programming and optimization experience (CUDA and/or Triton) — kernel-level tuning, memory hierarchy, and bandwidth optimization at scale.
  • Hands-on experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.
  • Hands-on experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.
  • Working knowledge of ML framework internals (PyTorch and/or JAX; torch.compile, XLA, or similar compiler paths).
  • Strong proficiency in Python, with the ability to drop into C++/CUDA (and Rust or Go) as the work demands.
  • High-ownership mindset — you measure yourself by throughput shipped and latency cut, not tickets closed.

Skills

  • CUDA
  • Triton
  • Python
  • C++
  • Rust
  • Go
  • PyTorch
  • JAX
  • torch.compile
  • XLA
  • Profiling
  • Roofline analysis
  • Latency optimization
  • Throughput optimization
  • Root-cause investigation
  • Kernel-level tuning
  • Memory hierarchy optimization
  • Bandwidth optimization
  • Inference optimization
  • Serving optimization
  • Batching
  • KV/prompt caching
  • Quantization
  • Low-latency sampling
  • High-throughput sampling
  • Training performance optimization
  • Parallelism
  • Distributed communication
  • Mixed precision
  • GPU utilization
  • Numerical correctness
  • FP8/INT8 quantization
  • Distributed systems
  • Collective communication (NCCL)
  • Interconnects (NVLink)
  • Model parallelism
  • Tensor parallelism
  • Fault tolerance
  • Generative models
  • Diffusion models
  • Video models
  • 3D/spatial models
  • TPU
  • Trainium
  • Performance modeling
  • Observability

Location

  • San Francisco Bay Area

Work Type

  • Onsite

Experience Level

  • Individual Contributor

Salary/Compensations

  • $200k-$300k base salary

Benefits

  • Equity awards

About the Company

  • World Labs is a frontier AI research and product company advancing spatial intelligence, the next frontier beyond large language models.
  • Co-founded by Dr. Fei-Fei Li, Justin Johnson and Ben Mildenhall, the company is pioneering world models that perceive, generate, reason, and interact with virtual and physical worlds.
  • The company’s flagship product, Marble, transforms text, images, and video into fully navigable 3D worlds, unlocking applications across gaming, film, architecture, robotics, and immersive digital experiences.
  • Backed by leading investors and with over $1B raised, World Labs is assembling a world-class team at the intersection of AI research and real-world deployment.

Equal Opportunity

  • World Labs is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected under applicable law. We welcome all qualified applicants and are committed to providing reasonable accommodations throughout the hiring process upon request.
  • In accordance with California law, we disclose the following: Pay Range $200-$300k base salary (good-faith estimate for San Francisco Bay Area upon hire; actual offer based on experience, skills, and qualifications). Total Compensation: Base salary plus equity awards. Salary History: We do not request or consider prior compensation in making offers. Compliance: Cal. Lab. Code §432.3 (pay scale disclosure & salary history ban); Cal. Lab. Code §1197.5 (Equal Pay Act); Cal. Gov. Code §12940 (FEHA); 42 U.S.C. §2000e (Title VII); 29 U.S.C. §621 (ADEA); 42 U.S.C. §12101 (ADA).