Research Engineer, Infrastructure, Training Systems at Thinking Machines Lab | CA, US | Rezi

Research Engineer, Infrastructure, Training Systems at Thinking Machines Lab

Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab · CA, US

3 weeks ago

Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking an infrastructure research engineer to design and build core systems for scalable, efficient training of large AI models. The goal is to ensure experimentation and training are fast and reliable, allowing research teams to focus on science.

Responsibilities

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes for large-scale training workloads.
  • Develop high-performance optimizations to maximize throughput and efficiency.
  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.
  • Establish standards for reliability, maintainability, and security, ensuring systems are robust under rapid iteration.
  • Collaborate with researchers and engineers to build scalable infrastructure.
  • Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
  • Past experience working on distributed training for the world’s largest models to make them stable, reliable, and performant.
  • Track record of improving research productivity through infrastructure design or process improvements.
  • Contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.

Skills

  • Distributed training systems
  • High-performance optimizations
  • Reusable frameworks and libraries
  • Deep learning frameworks (PyTorch, JAX)
  • System architectures
  • Reliability, maintainability, and security standards
  • Collaboration
  • Open-source contributions

Location

  • San Francisco, California

Work Type

  • Onsite

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.

Equal Opportunity

  • We sponsor visas.