Evals Engineer, Offensive Cyber at Zealot Labs | NY, US | Rezi

Evals Engineer, Offensive Cyber at Zealot Labs

Evals Engineer, Offensive Cyber

Zealot Labs · NY, US

1 weeks ago

Evals Engineer, Offensive Cyber

Zealot Labs · NY, US

9 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Evals Engineer, Offensive Cyber role.

Rezi rewrites your resume against Zealot Labs's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Evals Engineer, Offensive Cyber posting at Zealot Labs — free, in seconds.

About the Role

We are seeking an Evals Engineer to own the truth layer for offensive cyber, responsible for benchmarks, target environments, grading harnesses, instrumentation, metrics, and dashboards that determine the capabilities of our AI systems in vulnerability research, exploit development, and autonomous operations.

Responsibilities

  • Build CTF- and AIxCC-style task suites that reflect real targets and conditions.
  • Design eval methodology for non-deterministic agents, including variance across runs, sampling strategies, confidence thresholds, and statistically sound claims of capability.
  • Instrument full agent trajectories: tool calls, intermediate state, decision points, failed paths, partial progress, and final outcomes.
  • Build graders robust to reward hacking, including trustworthy LLM-as-judge pipelines and scoring for multi-stage exploitation chains.
  • Translate results into metrics and dashboards the research team uses to prioritize, and that hold up to scrutiny.

Requirements

  • Experience building LLM, model, or agent evaluation systems, with an understanding of contamination, benchmark overfit, grader drift, prompt sensitivity, brittle scoring, reward hacking, and false confidence.
  • Strong systems engineering ability; you can build reliable, reproducible evaluation infrastructure at scale.
  • A genuine security background.
  • Hands-on offensive experience in CTFs, vulnerability research, exploit development, reverse engineering, or AIxCC-style environments is a major advantage.
  • Rigor about measurement.

Skills

  • LLM evaluation
  • Model evaluation
  • Agent evaluation
  • Systems engineering
  • Offensive security
  • CTF
  • Vulnerability research
  • Exploit development
  • Reverse engineering
  • AIxCC
  • Measurement

About the Company

  • Our eval stack is the company's truth layer. It tells us whether our AI systems can actually find vulnerabilities, reason through exploit chains, operate autonomously, and improve in ways that are real rather than cosmetic.
  • If our evals are weak, we do not know what we have built. If they are strong, the entire research team moves faster with confidence.