About the Role
We are seeking an infrastructure research engineer to design and build core systems for scalable, efficient training of large AI models. The goal is to ensure experimentation and training are fast and reliable, allowing research teams to focus on science.
Responsibilities
- Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes for large-scale training workloads.
- Develop high-performance optimizations to maximize throughput and efficiency.
- Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.
- Establish standards for reliability, maintainability, and security, ensuring systems are robust under rapid iteration.
- Collaborate with researchers and engineers to build scalable infrastructure.
- Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
- A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
- Past experience working on distributed training for the world’s largest models to make them stable, reliable, and performant.
- Track record of improving research productivity through infrastructure design or process improvements.
- Contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.
Skills
- Distributed training systems
- High-performance optimizations
- Reusable frameworks and libraries
- Deep learning frameworks (PyTorch, JAX)
- System architectures
- Reliability, maintainability, and security standards
- Collaboration
- Open-source contributions
Location
- San Francisco, California
Work Type
- Onsite
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree or equivalent experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
Equal Opportunity
- We sponsor visas.
