Product Manager, APEX at Mercor | CA, US | Rezi

Product Manager, APEX at Mercor

Product Manager, APEX

Mercor · CA, US

1 months ago

Product Manager, APEX

Mercor · CA, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Product Manager, APEX role.

Rezi rewrites your resume against Mercor's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Product Manager, APEX posting at Mercor — free, in seconds.

About the Role

Mercor's AI Productivity Index (APEX) assesses how effectively frontier AI models can perform economically valuable work. We're looking for a Product Manager to own and scale the APEX brand and public leaderboard. In this role, you will define the positioning, strategy, roadmap, and operational excellence of Mercor's evaluation products, benchmarks, and public and private leaderboards. You will work across Research, Engineering, Operations, and Go-To-Market teams to transform evaluation datasets into trusted industry benchmarks that influence model development and purchasing decisions across the AI ecosystem. You will serve as the product owner for APEX and our eval platform, driving benchmark innovation, evaluation integrity, infrastructure, customer adoption, and business impact. You will work directly with frontier AI labs and enterprise customers, representing Mercor as a thought leader in AI evaluation and measurement. You will partner with Mercor’s world-class benchmark research team, which includes the first authors from many popular benchmarks including Tau Bench, SciCode, PostTrainBench, and more. The ideal candidate combines strong product judgment, technical fluency, operational rigor, and customer-facing experience, with a passion for turning emerging model capabilities into a credible measure of what AI can actually do for the economy.

Responsibilities

  • Own the roadmap and portfolio strategy, setting and influencing priorities across new benchmark development, leaderboard launches, infrastructure investment, and expansion into new evaluation categories.
  • Run the intake for new benchmarks, evaluating and prioritizing proposals from research, customers, and GTM against real demand and company strategy.
  • Enforce eval integrity, owning contamination policy, holdout strategy, versioning, auditability, and release cadence.
  • Build end-to-end eval pipeline, owning the path from eval run to published result.
  • Work directly with labs and customers to understand evaluation needs, present and defend results, and translate feedback into the next benchmark.
  • Close the loop to the business by connecting leaderboard demand signals to loss analysis investments and dataset production.
  • Get in the weeds by writing specs and PRDs, reading trajectories, spot-checking failures, and making small PRs to unblock yourself and continuously improve the system.

Requirements

  • 5+ years in product management, technical program management, or a customer-facing technical role.
  • Prior background in SWE, ML, or DS strongly preferred.
  • Fluent reasoning about rubric design, inter-rater reliability, agentic harnesses, contamination and overfitting, and what a small sample can and can't support.
  • Ability to distinguish a real capability gap from measurement noise.
  • Ability to be neutral, precise, and defensible under scrutiny when publishing numbers about other people's models.
  • Strong instincts for which capabilities are actually worth measuring, and what results are actually worth highlighting.
  • Track record of building a product, program, or business line from nothing.
  • Take full accountability for outcomes, not just outputs.
  • Able to self-direct in ambiguous contexts, creating clarity for others.
  • Ability to bring research, engineering, operations, and GTM around a shared methodology.

Skills

  • Product Management
  • Technical Program Management
  • Customer-facing technical role
  • SWE
  • ML
  • DS
  • Rubric design
  • Inter-rater reliability
  • Agentic harnesses
  • Contamination
  • Overfitting
  • Public judgment
  • Stakeholder alignment

Location

  • San Francisco
  • NYC
  • London

Work Type

  • In-person five days a week

Experience Level

  • 5+ years

Benefits

  • Bi-annual performance bonus structure
  • Generous equity grant vested over 4 years
  • Up to $15k Relocation bonus
  • $10K housing bonus (if you live within 0.5 miles of our office)
  • $1.5K monthly stipend for meals
  • Free Equinox membership
  • $200 monthly laundry reimbursement
  • $200 monthly personal wellness reimbursement
  • Health, Dental, Vision insurance

About the Company

  • Mercor's mission is to organize human intelligence to power the AI economy.
  • We're a leading AI data company, building the layer between human expertise and frontier models.
  • Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models.
  • Mercor's APEX benchmark family measures AI's real-world impact on professional work.
  • Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
  • Mercor is creating a new category of work where expertise powers AI advancement.
  • Achieving this requires an ambitious, fast-paced and deeply committed team.
  • You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society.
  • Mercor is a profitable Series C company valued at $10 billion.