Research Engineer, Infrastructure, RL Systems at Thinking Machines Lab | CA, US | Rezi

Research Engineer, Infrastructure, RL Systems at Thinking Machines Lab

Research Engineer, Infrastructure, RL Systems

Thinking Machines Lab · CA, US

3 weeks ago

Research Engineer, Infrastructure, RL Systems

Thinking Machines Lab · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking an infrastructure research engineer to design and build core systems for scalable, efficient training of large models using reinforcement learning. This role combines research and large-scale systems engineering, requiring expertise in RL algorithms and distributed training realities. You will optimize pipelines, enhance reliability and observability, and collaborate with researchers and infra teams to make reinforcement learning production-ready.

Responsibilities

  • Design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads.
  • Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
  • Develop shared monitoring and observability tools for RL systems.
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines.
  • Build evaluation and benchmarking infrastructure for model helpfulness, safety, and factuality.
  • Publish and share learnings through internal documentation, open-source libraries, or technical reports.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • Ability to thrive in a highly collaborative environment with cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative and work across different stacks and teams.
  • Experience training or supporting large-scale language models with tens of billions of parameters or more.
  • Experience working with reinforcement learning workloads (e.g., PPO, DPO, RLHF, or reward modeling).
  • Background in high-performance or reliability engineering — distributed training frameworks and cluster orchestration (Kubernetes, Slurm).
  • Familiarity with monitoring and observability tools (Prometheus, Grafana, OpenTelemetry).
  • Contributions to large-scale ML research or infrastructure, open-source frameworks, or internal performance optimization efforts.

Skills

  • Reinforcement Learning
  • Large-scale systems engineering
  • Distributed training
  • Inference at scale
  • Rollout optimization
  • Reward pipeline optimization
  • Reliability engineering
  • Observability
  • Orchestration
  • Deep learning frameworks (PyTorch, JAX)
  • High-performance computing
  • Kubernetes
  • Slurm
  • Prometheus
  • Grafana
  • OpenTelemetry

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.

Equal Opportunity

  • We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.