About the Role
As a safety researcher, you will work towards ensuring our models are safe and trustworthy, sitting at the intersection of research and hands-on technical work. You will explore the science behind how models handle harmful or dual-use requests and design experiments to inform model training and evaluation.
Responsibilities
- Build data filtering pipelines and quality classifiers to shape what models learn from pre-training corpora, and study how those early interventions affect downstream safety behavior.
- Apply post-training techniques, including RL from human and AI feedback and policy-based reasoning approaches, to shape how models handle harmful, sensitive, and dual-use requests.
- Design, build, and maintain safety evaluations, with particular focus on measuring model behavior on long-horizon and agentic tasks.
- Generate and curate synthetic data to train and evaluate models on refusal boundaries and safety-relevant behaviors.
- Red-team models and products to surface failure modes, jailbreaks, and emergent risks before deployment, and design mitigations for what is found.
Requirements
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- Background in AI safety research, with hands-on experience in at least one area of safety, such as: RLHF/RLAIF, alignment and preference modeling, deliberative alignment, safety evaluations, or red-teaming.
- Proficiency in Python and familiarity with deep learning frameworks (e.g., PyTorch, TensorFlow, or JAX).
- Comfort debugging distributed training and writing code that scales.
- Clarity in communication, an ability to explain complex technical concepts in writing.
- Experience building evaluations for long-horizon, multi-step, or agentic tasks.
- Experience generating synthetic data at scale for training or evaluation.
- Experience with modern red-teaming/jailbreaking techniques.
- Research contributions in AI safety — publications, open-source evaluations, or public red-teaming work.
- Familiarity with the AI safety literature and current open problems (e.g., scalable oversight, reward hacking, jailbreak robustness).
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
Skills
- Python
- Deep learning frameworks (PyTorch, TensorFlow, JAX)
- AI safety research
- RLHF/RLAIF
- Alignment and preference modeling
- Deliberative alignment
- Safety evaluations
- Red-teaming
- Communication
- Technical writing
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree
- PhD
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Health benefits
- Dental benefits
- Vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Visa sponsorship
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
Equal Opportunity
- As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
