Research Engineer, Infrastructure, Numerics at Thinking Machines Lab | CA, US | Rezi

Research Engineer, Infrastructure, Numerics at Thinking Machines Lab

Research Engineer, Infrastructure, Numerics

Thinking Machines Lab · CA, US

3 weeks ago

Research Engineer, Infrastructure, Numerics

Thinking Machines Lab · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking an infrastructure research engineer to design and build core systems for efficient large-scale model training, with a focus on numerics. You will improve the numerical foundations of our distributed training stack, enhancing stability, scalability, and speed for trillion-parameter models. This role is for individuals who excel at the intersection of research and systems engineering, understanding both optimization mathematics and distributed compute realities.

Responsibilities

  • Design and optimize distributed training infrastructure for large-scale LLMs, focusing on performance, stability, and reproducibility across multi-GPU and multi-node setups.
  • Implement and evaluate low-precision numerics (e.g., BF16, MXFP8, NVFP4) to improve efficiency without sacrificing model quality.
  • Develop kernels and communication primitives that utilize hardware-level support for mixed and low-precision arithmetic.
  • Collaborate with research teams to co-design model architectures and training recipes aligned with emerging numeric formats and stability constraints.
  • Prototype and benchmark scaling strategies such as data, tensor, and pipeline parallelism that integrate precision-adaptive computation and quantized communication.
  • Contribute to the design of internal orchestration and monitoring systems for efficient and reproducible distributed experiments.
  • Publish and share learnings through internal documentation, open-source libraries, or technical reports to advance scalable AI infrastructure.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
  • Ability to thrive in a highly collaborative environment with cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative and work across different stacks and teams.
  • Strong engineering skills, ability to contribute performant, maintainable code and debug complex codebases in areas such as floating-point numerics, low-precision arithmetic, and distributed systems.
  • Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM.
  • Experience implementing FP8, INT8, or block-floating point (MX) formats and understanding their numerical trade-offs.
  • Prior contributions to open-source deep learning infrastructure such as PyTorch, DeepSpeed, or XLA.
  • Publications, patents, or projects related to numerical optimization, communication-efficient training, or systems for large models.
  • Experience training and supporting large-scale AI models.
  • Track record of improving research productivity through infrastructure design or process improvements.

Skills

  • Deep learning frameworks (PyTorch, JAX)
  • Distributed systems
  • Floating-point numerics
  • Low-precision arithmetic
  • Numerical optimization
  • Communication-efficient training
  • Systems for large models
  • PyTorch/XLA
  • DeepSpeed
  • Megatron-LM
  • FP8
  • INT8
  • Block-floating point (MX)

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.