About the Role
We are seeking an infrastructure research engineer to design and build core systems for efficient large-scale model training, with a focus on numerics. You will improve the numerical foundations of our distributed training stack, enhancing stability, scalability, and speed for trillion-parameter models. This role is for individuals who excel at the intersection of research and systems engineering, understanding both optimization mathematics and distributed compute realities.
Responsibilities
- Design and optimize distributed training infrastructure for large-scale LLMs, focusing on performance, stability, and reproducibility across multi-GPU and multi-node setups.
- Implement and evaluate low-precision numerics (e.g., BF16, MXFP8, NVFP4) to improve efficiency without sacrificing model quality.
- Develop kernels and communication primitives that utilize hardware-level support for mixed and low-precision arithmetic.
- Collaborate with research teams to co-design model architectures and training recipes aligned with emerging numeric formats and stability constraints.
- Prototype and benchmark scaling strategies such as data, tensor, and pipeline parallelism that integrate precision-adaptive computation and quantized communication.
- Contribute to the design of internal orchestration and monitoring systems for efficient and reproducible distributed experiments.
- Publish and share learnings through internal documentation, open-source libraries, or technical reports to advance scalable AI infrastructure.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
- Ability to thrive in a highly collaborative environment with cross-functional partners and subject matter experts.
- A bias for action with a mindset to take initiative and work across different stacks and teams.
- Strong engineering skills, ability to contribute performant, maintainable code and debug complex codebases in areas such as floating-point numerics, low-precision arithmetic, and distributed systems.
- Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM.
- Experience implementing FP8, INT8, or block-floating point (MX) formats and understanding their numerical trade-offs.
- Prior contributions to open-source deep learning infrastructure such as PyTorch, DeepSpeed, or XLA.
- Publications, patents, or projects related to numerical optimization, communication-efficient training, or systems for large models.
- Experience training and supporting large-scale AI models.
- Track record of improving research productivity through infrastructure design or process improvements.
Skills
- Deep learning frameworks (PyTorch, JAX)
- Distributed systems
- Floating-point numerics
- Low-precision arithmetic
- Numerical optimization
- Communication-efficient training
- Systems for large models
- PyTorch/XLA
- DeepSpeed
- Megatron-LM
- FP8
- INT8
- Block-floating point (MX)
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree or equivalent experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Visa sponsorship
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
