Member of Technical Staff, Enterprise Evals Platform at Mercor | CA, US | Rezi

Member of Technical Staff, Enterprise Evals Platform at Mercor

Member of Technical Staff, Enterprise Evals Platform

Mercor · CA, US

1 months ago

Member of Technical Staff, Enterprise Evals Platform

Mercor · CA, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Member of Technical Staff, Enterprise Evals Platform role.

Rezi rewrites your resume against Mercor's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Member of Technical Staff, Enterprise Evals Platform posting at Mercor — free, in seconds.

About the Role

You will apply Mercor's learnings from building benchmarks with domain experts to devise new methods for improving evals, rubrics, and the agents measured against them. This platform engineering role requires a deep understanding of evals to build verifiers, agent measurement environments, and scalable grading infrastructure.

Responsibilities

  • Define golden sets by decomposing real tasks and encoding the expert quality bar.
  • Build verifiers over agent trajectories and outputs that are calibrated and difficult to game.
  • Construct the eval platform for offline environments, task suites, and scalable grading.
  • Conduct loss analysis on production trajectories and convert failure modes into regression tests.
  • Manage the optimization loop across models, prompts, skills, and harnesses.
  • Oversee rollout gates for agent changes.
  • Collaborate with the Enterprise Platform team and customer-embedded Applied AI engineers.

Requirements

  • Professional, academic, or research experience in agent engineering and evaluation, including agent runtimes, harnesses, trajectories, and failure points.
  • Experience building evaluation suites for LLM or agent systems.
  • Familiarity with how benchmarks like terminal-bench, tau-bench, and APEX are constructed and exploited.
  • Judgment in task and rubric design, translating fuzzy quality notions into measurable metrics with demonstrable agent or model improvements.
  • Strong software engineering fundamentals.
  • Ability to work independently on ambiguous, loosely specified problems.

Skills

  • Agent engineering
  • Agent evaluation
  • LLM evaluation
  • Software engineering
  • Harbor environments
  • RL environments

Location

  • San Francisco
  • NYC
  • London

Work Type

  • In-person
  • Full-time

Experience Level

  • Professional
  • Academic
  • Research

Benefits

  • Up to $15k relocation bonus
  • $10K housing bonus (if living within 0.5 miles of office)
  • $1.5K monthly stipend for meals
  • Generous equity grant vested over 4 years
  • Free Equinox membership
  • $200 monthly laundry reimbursement
  • $200 monthly personal wellness reimbursement
  • Health, Dental, Vision insurance

About the Company

  • Mercor is a leading AI data company organizing human intelligence to power the AI economy.
  • We build the layer between human expertise and frontier models.
  • Millions of domain experts are paid daily to train frontier AI models.
  • Mercor's APEX benchmark family measures AI's real-world impact on professional work.
  • Mercor Enterprise provides infrastructure for Fortune 500 companies to capture and translate employee expertise into agents.
  • We are creating a new category of work where expertise powers AI advancement.
  • Mercor is a profitable Series C company valued at $10 billion.
  • We work in-person five days a week in our San Francisco, NYC, or London offices.