About the Role
This research role focuses on scaling reinforcement learning for frontier models, with a core interest in asynchronous RL and co-designing the RL recipe with its supporting systems. The position requires a deep understanding of both ML and systems aspects, particularly concerning inference efficiency and large-scale training.
Responsibilities
- Co-design the RL recipe and the systems that run it, making joint recipe-level and systems-level choices and validating them at frontier scale.
- Advance asynchronous RL algorithms.
- Improve the efficiency of rollout generation and its integration with training, treating inference as a first-class part of the RL loop.
- Run frontier-scale RL end to end, including bringing up new models and training setups, and maintaining stability of large runs.
- Jointly optimize the compute and training efficiency of RL, focusing on accelerator utilization, memory, communication, and low-precision numerics.
- Conduct empirical science, including ablations and scaling studies, supported by reliable instrumentation and clear documentation.
Requirements
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
- Comfortable with debugging distributed training and writing code that scales.
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- Clarity in communication and the ability to explain complex technical concepts in writing.
- Strong research judgment, demonstrated through clean ablations, honest baselines, and clear technical writing.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
- Strong grounding in RL for large language models, such as modern policy optimization methods and their behavior at scale.
- Deep understanding of asynchronous RL: the algorithms, design choices, and trade-offs, on both the ML and the systems sides.
- Experience training large models across many accelerators, with comfort inside the distributed stack (parallelism strategies, memory, communication).
- Good working knowledge of inference systems: able to reason quantitatively about rollout generation throughput and cost.
- Experience building or operating decoupled generation/training RL systems at scale.
- Experience with RL on verifiable and agentic tasks, including multi-turn environments.
- Experience with RL training stability techniques for large runs.
- Familiarity with low-precision training and inference: numerics, quantization, and their implications for RL.
- Hands-on work with LLM serving stacks (e.g., SGLang, vLLM, TokenSpeed, or custom engines).
- Experience with scaling studies for large models.
- Contributions to open-source training or inference frameworks.
Skills
- Python
- Deep learning frameworks (PyTorch, TensorFlow, JAX)
- Distributed training
- Reinforcement Learning (RL)
- Asynchronous RL
- Inference systems
- Low-precision numerics
- Quantization
- LLM serving stacks (SGLang, vLLM, TokenSpeed, or custom engines)
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Research
Education Level
- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline or equivalent industry research experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
Equal Opportunity
- As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
