Research, RL Scaling at Thinking Machines Lab | CA, US | Rezi

Research, RL Scaling at Thinking Machines Lab

Research, RL Scaling

Thinking Machines Lab · CA, US

1 weeks ago

Research, RL Scaling

Thinking Machines Lab · CA, US

8 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

This research role focuses on scaling reinforcement learning for frontier models, with a core interest in asynchronous RL and co-designing the RL recipe with its supporting systems. The position requires a deep understanding of both ML and systems aspects, particularly concerning inference efficiency and large-scale training.

Responsibilities

  • Co-design the RL recipe and the systems that run it, making joint recipe-level and systems-level choices and validating them at frontier scale.
  • Advance asynchronous RL algorithms.
  • Improve the efficiency of rollout generation and its integration with training, treating inference as a first-class part of the RL loop.
  • Run frontier-scale RL end to end, including bringing up new models and training setups, and maintaining stability of large runs.
  • Jointly optimize the compute and training efficiency of RL, focusing on accelerator utilization, memory, communication, and low-precision numerics.
  • Conduct empirical science, including ablations and scaling studies, supported by reliable instrumentation and clear documentation.

Requirements

  • Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
  • Comfortable with debugging distributed training and writing code that scales.
  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
  • Clarity in communication and the ability to explain complex technical concepts in writing.
  • Strong research judgment, demonstrated through clean ablations, honest baselines, and clear technical writing.
  • PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
  • Strong grounding in RL for large language models, such as modern policy optimization methods and their behavior at scale.
  • Deep understanding of asynchronous RL: the algorithms, design choices, and trade-offs, on both the ML and the systems sides.
  • Experience training large models across many accelerators, with comfort inside the distributed stack (parallelism strategies, memory, communication).
  • Good working knowledge of inference systems: able to reason quantitatively about rollout generation throughput and cost.
  • Experience building or operating decoupled generation/training RL systems at scale.
  • Experience with RL on verifiable and agentic tasks, including multi-turn environments.
  • Experience with RL training stability techniques for large runs.
  • Familiarity with low-precision training and inference: numerics, quantization, and their implications for RL.
  • Hands-on work with LLM serving stacks (e.g., SGLang, vLLM, TokenSpeed, or custom engines).
  • Experience with scaling studies for large models.
  • Contributions to open-source training or inference frameworks.

Skills

  • Python
  • Deep learning frameworks (PyTorch, TensorFlow, JAX)
  • Distributed training
  • Reinforcement Learning (RL)
  • Asynchronous RL
  • Inference systems
  • Low-precision numerics
  • Quantization
  • LLM serving stacks (SGLang, vLLM, TokenSpeed, or custom engines)

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • Research

Education Level

  • Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline
  • PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline or equivalent industry research experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.

Equal Opportunity

  • As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.