Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Evals Lead role.
Rezi rewrites your resume against Build AI's job description. Free.

Tailor your resume to this Evals Lead role.
Rezi rewrites your resume against Build AI's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Evals Lead posting at Build AI — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Evals Lead posting at Build AI — free, in seconds.
About the Role
We’re hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You need to deeply understand the philosophy of evals and have pushed consequential evals before.
Responsibilities
- Design and run benchmarks for video and world models, with Build data and without it, under the same protocol
- Make the comparison honest: held-out tasks, contamination and leakage checks, no cooking the numbers
- Build the analytics layer: model performance, failure patterns, and whether scaling the dataset moves capability — and on which axes
- Work with research and Head of Dataset & Quality so collection and evals inform each other
- Push evals that are consequential enough that people change plans when the number moves
- Design processes that increase evaluation quality, repeatability, and scale as we add tasks and countries
Requirements
- Shipped or driven evals that mattered: they changed training, hiring, product, or data decisions
- Understand eval philosophy well enough to argue about validity, not only to plot a curve
- The bar is consequential evals more than a specific domain
- Will not confuse a pretty dashboard with an eval that is allowed to decide things
Skills
- Video
- Robotics
- World models
- Multimodal evals
- Construct validity
- Contamination
- Leakage
- Evaluation operations: throughput, failure taxonomy, repeatability
Location
- San Francisco (Financial District)
- Shenzhen (Nanshan)
Work Type
- Fully in-person
Experience Level
- Lead
Benefits
- Competitive pay
- Medical, dental, and vision packages with generous premium coverage
- $500 per month credit for waiving medical benefits
- Housing subsidy of $2k per month for those living within walking distance of the office
- Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)
- Various wellness benefits covering fitness, mental health, and more
- Daily lunch and dinner in our office
- Unlimited compute budget subject to ROI justification
- Travel
About the Company
- Build AI is the data hyperscaler for Physical AI. We co-design hardware, collection, infrastructure, and research to scale the physical labor dataset orders of magnitude faster than anyone in the world.
- Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.
- We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Equal Opportunity
- Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply.