About the Role
We are seeking an infrastructure research engineer to design and build core systems for scalable, efficient training of large models using reinforcement learning. This role combines research and large-scale systems engineering, requiring expertise in RL algorithms and distributed training realities. You will optimize pipelines, enhance reliability and observability, and collaborate with researchers and infra teams to make reinforcement learning production-ready.
Responsibilities
- Design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads.
- Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
- Develop shared monitoring and observability tools for RL systems.
- Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines.
- Build evaluation and benchmarking infrastructure for model helpfulness, safety, and factuality.
- Publish and share learnings through internal documentation, open-source libraries, or technical reports.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
- Ability to thrive in a highly collaborative environment with cross-functional partners and subject matter experts.
- A bias for action with a mindset to take initiative and work across different stacks and teams.
- Experience training or supporting large-scale language models with tens of billions of parameters or more.
- Experience working with reinforcement learning workloads (e.g., PPO, DPO, RLHF, or reward modeling).
- Background in high-performance or reliability engineering — distributed training frameworks and cluster orchestration (Kubernetes, Slurm).
- Familiarity with monitoring and observability tools (Prometheus, Grafana, OpenTelemetry).
- Contributions to large-scale ML research or infrastructure, open-source frameworks, or internal performance optimization efforts.
Skills
- Reinforcement Learning
- Large-scale systems engineering
- Distributed training
- Inference at scale
- Rollout optimization
- Reward pipeline optimization
- Reliability engineering
- Observability
- Orchestration
- Deep learning frameworks (PyTorch, JAX)
- High-performance computing
- Kubernetes
- Slurm
- Prometheus
- Grafana
- OpenTelemetry
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree or equivalent experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
Equal Opportunity
- We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
