Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Research Lead, Evaluations and Benchmarks role.
Rezi rewrites your resume against Alice's job description. Free.

Tailor your resume to this Research Lead, Evaluations and Benchmarks role.
Rezi rewrites your resume against Alice's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Research Lead, Evaluations and Benchmarks posting at Alice — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Research Lead, Evaluations and Benchmarks posting at Alice — free, in seconds.
About the Role
You will ship a benchmark every two to three weeks, measuring frontier risks. Some benchmarks will be public, while others will be shared internally. You will collaborate with in-house researchers and direct freelancers, owning the taxonomy, harness, quality bar, and release process. This role is within the CTO office, working alongside the research lead and collaborating with approximately 150 researchers focused on AI harms.
Responsibilities
- Ship a benchmark every two to three weeks, with size varying by subject.
- Maintain a high-quality bar for benchmarks, ensuring reproducibility and novelty.
- Manage the evaluation process, including timelines and directing freelancers.
- Collaborate monthly with CTO and research leads to set the quarterly release plan based on research, client needs, and industry trends.
- Dedicate approximately 20% of time to staying current with AI safety and security research, maintaining lab contacts, and attending conferences.
Requirements
- PhD or Masters in computer science, machine learning, or a related field, or equivalent industry research experience.
- 3+ years of experience building and running safety or security evaluations for language models in production.
- 5+ relevant research publications in AI safety and security, with at least 2 as lead author.
- Strong engineering skills, including evaluation harnesses, distributed inference, and code analysis.
- Ability to develop taxonomies, not just score against them.
- Experience directing researchers and freelancers without formal management.
- Strong written and spoken English for effective internal communication across time zones.
- Curiosity and willingness to learn new subjects every three weeks.
- Post-training experience (SFT, DPO, GRPO) is ideal.
- Agentic evaluation experience (tool use, orchestration, permissions, prompt injection) is ideal.
- Publications at top conferences are ideal.
- Willingness to present work on client calls.
- Ability to present to large and/or senior audiences.
- Willingness to travel to conferences at least 3 times a year.
Skills
- Evaluation harnesses
- Distributed inference
- vLLM
- Codebase analysis and modification
- Taxonomy development
- Researcher and freelancer direction
- AI safety
- AI security
- Language model evaluation
- SFT
- DPO
- GRPO
- Reward design
- Agentic evaluation
- Tool use
- Orchestration
- Permissions
- Prompt injection
- Communication (written and spoken)
- Presentation skills
Location
- Remote
- Onsite
Work Type
- Full-time
Experience Level
- 3+ years
- 5+ years
Education Level
- PhD
- Masters
About the Company
- Alice is a trust, safety, and security company for the AI era, safeguarding communicative technologies.
- Alice serves the top 8 AI Labs globally, providing end-to-end AI lifecycle coverage.
- Solutions include model hardening evaluations, pre-deployment red-teaming, runtime guardrails, and drift detection.
- Alice is a recognized global leader in online safety and AI security.
- The company employs forward-thinking individuals dedicated to safeguarding over 3 billion users across major AI and tech platforms.