Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Site Reliability Engineer, Post Training role.
Rezi rewrites your resume against Thinking Machines Lab's job description. Free.

Tailor your resume to this Site Reliability Engineer, Post Training role.
Rezi rewrites your resume against Thinking Machines Lab's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Site Reliability Engineer, Post Training posting at Thinking Machines Lab — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Site Reliability Engineer, Post Training posting at Thinking Machines Lab — free, in seconds.
About the Role
We are hiring a Site Reliability Engineer (SRE) to ensure our post-training and reinforcement learning (RL) systems are fast, reliable, and facilitate researcher iteration. This role involves owning the health of training runs, clusters, and pipelines, working closely with research teams to debug failures, harden infrastructure, and build automation.
Responsibilities
- Own the reliability, performance, and uptime of large-scale post-training and RL training jobs.
- Partner with research teams during active model runs to unblock training and speed up iteration.
- Debug failures across the full stack, including accelerators, networking, storage, schedulers, and training frameworks.
- Build monitoring, alerting, and automated recovery systems for training runs.
- Improve checkpointing, fault tolerance, and job scheduling to minimize compute loss from hardware failures.
- Build internal tools to reduce toil and enhance cluster utilization for post-training and RL workloads.
- Participate in an on-call rotation supporting production model runs.
- Write postmortems and implement permanent infrastructure fixes for recurring failure patterns.
Requirements
- 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production.
- Track record debugging complex failures across distributed systems (networking, hardware, kernel, or scheduler issues).
- Strong software engineering skills in Python and/or Go/C++.
- Solid grounding in Linux systems internals and networking fundamentals.
- Comfortable owning production systems, including participating in on-call rotations.
- Experience operating GPU or TPU training clusters at scale.
- Familiarity with post-training and RL techniques and their infrastructure challenges.
- Experience with distributed training frameworks and job schedulers.
- Experience with high-performance networking and its role in distributed training performance.
- Experience building observability tooling purpose-built for ML training.
- A track record of thriving in fast-changing, research-driven environments.
Skills
- Python
- Go/C++
- Linux systems internals
- Networking fundamentals
- GPU/TPU training clusters
- Post-training and RL techniques
- Distributed training frameworks (PyTorch, Ray)
- Job schedulers (Slurm, Kubernetes)
- High-performance networking (InfiniBand, RDMA, NCCL)
- ML training observability tooling
Location
- San Francisco, CA
Work Type
- Onsite
Experience Level
- 4+ years
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
- We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication.
- We believe the future worth building is human, and we're hiring people who want to build it.