About the Role
We are seeking an infrastructure research engineer to design, optimize, and scale systems for large AI models, aiming to make inference faster, more cost-effective, reliable, and reproducible. This role supports research and real-world applications by ensuring smooth, large-scale operation of experiments, evaluations, and deployments.
Responsibilities
- Work alongside researchers and engineers to bring cutting-edge AI models into production.
- Collaborate with research teams to enable high-performance inference for novel architectures.
- Design and implement new techniques, tools, and architectures that improve performance, latency, throughput, and efficiency.
- Optimize codebase and compute fleet (e.g., GPUs) to fully utilize hardware FLOPs, bandwidth, and memory.
- Extend orchestration frameworks (e.g., Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch serving.
- Establish standards for reliability, observability, and reproducibility across the inference stack.
- Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
- Experience with inference serving systems optimized for throughput and latency (e.g., SGLang, vLLM).
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
- A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases.
- Experience training or supporting large-scale language models with hundreds of billions of parameters or more.
- Understanding of distributed compute systems, GPU parallelism, and hardware-aware optimizations.
- Contributions to open-source ML or systems infrastructure projects (e.g., SGLang, vLLM, PyTorch, Triton, DeepSpeed, XLA).
- Track record of improving research productivity through infrastructure design or process improvements.
Skills
- Deep learning frameworks (e.g., PyTorch, JAX)
- Inference serving systems optimized for throughput and latency (e.g., SGLang, vLLM)
- Distributed compute systems
- GPU parallelism
- Hardware-aware optimizations
- ML or systems infrastructure projects (e.g., SGLang, vLLM, PyTorch, Triton, DeepSpeed, XLA)
Location
- San Francisco, California
Work Type
- Onsite
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree or equivalent experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support as needed
- Visa sponsorship
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
