About the Role
We are hiring a software engineer to build and scale the distributed systems powering Tinker, our post-training platform. You will work on the infrastructure for scheduling jobs across large GPU clusters, maintaining state consistency under failure, and providing a reliable path from idea to a customized model. This role is central to Tinker's reliability and scale, involving system design, operation, and incident response with significant ownership and minimal oversight.
Responsibilities
- Design, build, and operate the distributed systems underlying Tinker, including job scheduling, resource orchestration, checkpointing, and fault recovery across large multi-node GPU clusters.
- Own the reliability, latency, and throughput of live production systems serving concurrent training and sampling workloads for external users.
- Build for graceful degradation and fast recovery: handle node failures, preemptions, and network partitions without losing user progress or data.
- Improve observability across the platform — metrics, tracing, and alerting that make it possible to detect and diagnose issues before or as they affect users.
- Lead incident response for the systems you own, drive root-cause analysis, and turn findings into lasting fixes.
- Partner with research and ML infrastructure teams to understand new training and sampling workloads and to evolve the platform's architecture to support them at scale.
- Make build-vs-buy and architectural tradeoff calls for core infrastructure components, and mentor other engineers on distributed systems practices.
Requirements
- 7+ years of experience building, running, and scaling distributed systems in production, with direct on-call ownership of live services.
- Deep understanding of distributed systems fundamentals: consensus, consistency models, failure detection, scheduling, and state management under partial failure.
- Track record of designing systems that operate reliably at scale, including diagnosing and resolving production incidents under time pressure.
- Strong systems-level coding ability in a language such as Python, Go, Rust, or C++, and comfort working close to the infrastructure layer.
- Experience operating large-scale GPU or accelerator clusters, including job orchestration, resource scheduling, or ML training infrastructure.
- Familiarity with the training or inference stack for large language models, and an understanding of the systems demands unique to fine-tuning or post-training workloads.
- Experience with checkpointing, distributed storage, or high-throughput networking (e.g., RDMA, NCCL) in performance-critical settings.
- Experience building developer-facing infrastructure or APIs where reliability and usability are both first-class concerns.
- A track record of driving platform-wide architectural decisions and comfort operating with significant autonomy in a fast-changing environment.
Skills
- Python
- Go
- Rust
- C++
- Distributed Systems
- Job Scheduling
- Resource Orchestration
- Checkpointing
- Fault Recovery
- Reliability
- Latency
- Throughput
- Observability
- Incident Response
- Root-Cause Analysis
- ML Infrastructure
- GPU Clusters
- Large Language Models
- Fine-tuning
- Distributed Storage
- High-throughput Networking
- RDMA
- NCCL
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- 7+ years
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
- We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication.
- We believe the future worth building is human, and we're hiring people who want to build it.
- Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs — opening access to capabilities that have previously been concentrated in a handful of labs.
- We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs.
- Tinker is rapidly adding new customers, features, and novel use-cases, and the systems underneath it run as a live, always-on production platform.
