Research Scientist, APEX Benchmarks at Mercor | CA, US | Rezi

Research Scientist, APEX Benchmarks at Mercor

Research Scientist, APEX Benchmarks

Mercor · CA, US

3 weeks ago

Research Scientist, APEX Benchmarks

Mercor · CA, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Research Scientist, APEX Benchmarks role.

Rezi rewrites your resume against Mercor's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Research Scientist, APEX Benchmarks posting at Mercor — free, in seconds.

About the Role

We are seeking a Research Scientist to lead the design of the next generation of APEX benchmarks and the expert-built datasets that support them. This role is highly visible, operating at the intersection of research, company strategy, and go-to-market. You will determine future measurement priorities based on frontier model performance gaps, design benchmarks and scoring systems, collaborate with partners to build them, and work with data operations, product, and GTM teams to scale production. You will also serve as a credible external technical voice.

Responsibilities

  • Benchmark design: Determine future APEX benchmark focus areas based on frontier model performance, saturation of existing evaluations, and the accessibility of economically valuable work. Own task taxonomy, difficulty calibration, contamination controls, and statistical design.
  • Dataset design: Design expert-built datasets and grading rubrics at scale, defining task difficulty, defensible grading criteria, and quality assurance for parallel expert contributions.
  • Measurement rigor: Establish standards for reporting results, including confidence intervals, inter-rater agreement, human vs. model-as-judge calibration, held-out splits, and failure analysis.
  • Partnerships: Collaborate with academic and industry partners to co-design benchmarks and drive adoption.
  • External research voice: Publish research findings (arXiv papers, open datasets, blog posts, conference talks, leaderboard releases) and represent Mercor's research to frontier labs, customers, and the press.
  • Translate results into narrative: Convert benchmark findings into clear arguments about the ROI of expert-curated data for technical reports, customer discussions, and go-to-market materials.
  • Cross-functional work: Partner with data operations, engineering, product, and strategy teams to move benchmarks from design to production and to inform the company roadmap with research insights.
  • Stay at the frontier: Monitor LLM evaluation literature and integrate relevant advancements into Mercor's benchmark development processes.

Requirements

  • Strong applied or academic research background in LLM evaluation, benchmarking, NLP, or a related field, with a proven track record of rigorous experimental design.
  • Ability to identify informative measurements for frontier models based on their behavior, rather than solely on ease of implementation.
  • Proficiency in statistical reasoning regarding sampling, variance, contamination, and grader reliability.
  • Hands-on coding skills for building evaluation harnesses, running experiments, and analyzing results.
  • Exceptional communication skills for presenting complex technical findings clearly to both technical and non-technical audiences, in writing and in person.
  • Comfort operating in fast-moving, cross-functional environments with undefined problem spaces.
  • Genuine curiosity about go-to-market strategy, startup dynamics, and the business of AI data.
  • Excited to work in-person five days a week in our San Francisco office in a high-intensity, high-ownership environment.

Skills

  • LLM evaluation
  • Benchmarking
  • NLP
  • Experimental design
  • Statistical rigor
  • Coding
  • Communication
  • Cross-functional collaboration

Location

  • San Francisco
  • NYC
  • London

Work Type

  • In-person
  • Full-time

Experience Level

  • Research Scientist

Education Level

  • Ph.D. in machine learning, NLP or a related field; equivalent industry or frontier lab research experience considered.
  • Publications at top-tier venues (NeurIPS, ICML, ACL, ICLR), especially in evaluation, benchmarking or data-centric AI.
  • Experience authoring a widely-adopted public benchmark or dataset.
  • Industry experience on an evaluation, benchmarking or post-training team at a frontier lab.
  • Domain depth in one of APEX’s professional verticals — finance, law, consulting, accounting, medicine or software engineering.
  • Experience designing rubrics, model-as-judge pipelines, or human annotation programs at scale.

Benefits

  • Bi-annual performance bonus structure
  • Generous equity grant vested over 4 years
  • Up to $15k Relocation bonus
  • $10K housing bonus (if you live within 0.5 miles of our office)
  • $1.5K monthly stipend for meals
  • Free Equinox membership
  • $200 monthly laundry reimbursement
  • $200 monthly personal wellness reimbursement
  • Health, Dental, Vision insurance
  • 401(k) with company match

About the Company

  • Mercor is a leading AI data company organizing human intelligence to power the AI economy, building the layer between human expertise and frontier models.
  • Millions of domain experts are paid over $4 million per day to train frontier AI models.
  • Mercor's APEX benchmark family measures AI's real-world impact on professional work.
  • Mercor Enterprise brings this infrastructure to Fortune 500 companies to capture how their best people work and translate that expertise into agents.
  • Mercor is creating a new category of work where expertise powers AI advancement, requiring an ambitious, fast-paced, and deeply committed team.
  • The team works alongside researchers, operators, and AI companies at the forefront of shaping systems that are redefining society.
  • Mercor is a profitable Series C company valued at $10 billion.