About the Role
The Coding Agents team focuses on enhancing AI's ability to perform agentic coding tasks, including writing, debugging, and reasoning about code over extended, multi-turn interactions. You will join a small, high-impact team responsible for the methodologies, data, and infrastructure that drive coding capability improvements in model releases. This team manages the entire post-training coding process, encompassing synthetic and human data generation, RL environments, reward design, and large-scale training.
Responsibilities
- Design and execute RL training jobs focused on agentic coding capabilities, refining methodologies and data.
- Develop and enhance the sandboxed coding environments and reward mechanisms used for model training and evaluation.
- Generate and curate high-quality synthetic coding data, and construct scalable, general-purpose data pipelines.
- Design evaluations to measure real-world coding utility and train models against them to achieve tangible improvements in daily usability.
- Analyze large RL runs to identify and address issues such as confounders, reward hacking, and other RL failure modes.
- Collaborate with infrastructure, evaluation, and other post-training teams on shared data, joint training runs, and usability enhancements, integrating results into model releases.
Requirements
- Strong engineering skills with the ability to contribute code and debug complex codebases.
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
- Comfortable with debugging distributed training and writing scalable code.
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- Clear communication skills, with the ability to explain complex technical concepts in writing.
- Experience building synthetic data pipelines and systems that were adopted by others and remain in use.
- Experience managing the end-to-end process of identifying model usability gaps and addressing them through custom evaluations and training data.
- Experience ensuring the reliability of large-scale agentic RL infrastructure, accounting for long-tail failures.
- Experience improving the coding capabilities of a frontier model.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
Skills
- Python
- Deep learning frameworks (PyTorch, TensorFlow, JAX)
- Distributed training
- Scalable code development
- Synthetic data pipeline development
- Agentic RL infrastructure
- Model usability evaluation
- Communication
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Research role with ownership
- Industry research experience
Education Level
- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
- We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication.
- We believe the future worth building is human, and we're hiring people who want to build it.
Equal Opportunity
- We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
