About the Role
The Evaluation Data Scientist partners closely with the standards and evaluations team, owning the technical core of designing evaluation methodology, establishing statistical standards for defensible findings, and building infrastructure to maintain the Youth AI Safety Institute's scientific credibility. This role also serves as a senior data science partner to the broader organization, collaborating with editorial, education, research, and advocacy teams to build KPI frameworks, design impact analyses, and translate program goals into measurable outcomes.
Responsibilities
- Translate safety standards and SME judgment into measurable, reproducible evaluation protocols.
- Establish and defend statistical standards, including inter-rater reliability, agreement analysis, sample size and power calculations, and uncertainty quantification.
- Design and maintain the evaluation architecture, including simulated persona generation, multi-turn interaction harnesses, LLM-as-judge ensembles, and human-in-the-loop validation.
- Build the feedback loop to detect model drift and ensure evaluation comparability over time.
- Build internal tooling and dashboards for the Risk Assessment team and subject matter experts.
- Author methodology findings in AI evaluation venues and contribute technical rigor to research outputs.
- Represent the Institute's methods to academic reviewers, policymakers, and AI labs.
- Communicate findings to non-technical audiences, including executive leadership and government offices.
- Provide methodological oversight for external evaluation partners, reviewing their approaches and ensuring coherence.
- Ensure external partner outputs align with Common Sense Media's research and risk frameworks.
- Design and execute A/B tests, randomized controlled trials, and quasi-experimental studies.
- Lead rigorous program evaluations measuring the impact of Common Sense Media's work.
- Develop predictive and statistical models supporting organization-wide decisions.
- Partner with the data team on shared data infrastructure and methodology.
- Advise on methodology, raise the technical bar, and help non-technical colleagues use data and AI tools effectively.
Requirements
- 6–8+ years of applied data science and ML experience, with a recent focus on generative AI evaluation.
- Demonstrated command of the AI evaluation stack: LLM-as-judge design, synthetic persona generation, multi-turn testing, human-in-the-loop validation, and agreement statistics.
- Production fluency in Python, the modern ML/AI stack (transformers, LLM APIs, PyTorch/sklearn), and SQL.
- Experience designing evaluation instruments, ideally for youth-specific or sensitive-population risks.
- Ability to operate as a technical peer to senior engineers: code review, contributing to builds, shared architectural ownership.
- Experience applying causal inference and evaluation methods to measure real-world program or product outcomes.
- Strong written and verbal communication skills.
- Strong understanding of data governance, privacy, and responsible data practices, with sensitivity to data involving children and families.
Skills
- Generative AI evaluation
- LLM-as-judge design
- Synthetic persona generation
- Multi-turn testing
- Human-in-the-loop validation
- Agreement statistics
- Python
- ML/AI stack (transformers, LLM APIs, PyTorch/sklearn)
- SQL
- Causal inference
- Evaluation methods
- Data governance
- Privacy
- Responsible data practices
Location
- San Francisco, CA
Work Type
- Full-time
Experience Level
- 6-8+ years
Salary/Compensations
- $140,000.00 - $166,250.00
Benefits
- Competitive nonprofit compensation
- Comprehensive benefits
- Collaborative, flexible work environment with meaningful leadership access and visibility
About the Company
- Common Sense Media is the leading nonprofit organization dedicated to improving the lives of kids and families by providing research-backed information, education, and an independent voice. We rate, educate, and advocate for policies to protect and prepare kids online. Our ratings, research, and resources reach over 150 million users globally, 1.4 million educators, and over 100,000 schools worldwide annually.
- The Youth AI Safety Institute sets standards, conducts research, and independently tests AI products children use.
Equal Opportunity
- Common Sense Media provides equal employment opportunities to all qualified individuals and prohibits discrimination and harassment of any type without regard to race, color, religion, sex, gender identity, sexual orientation, pregnancy, age, national origin, physical or mental disability, military or veteran status, genetic information, or any other protected classification or characteristic protected by federal, state, or local laws.
- Common Sense Media will also consider for employment qualified applicants with arrest and conviction records.
- Pursuant to the San Francisco Fair Chance Ordinance, we will consider employment for qualified applicants with arrest and conviction records.
