Software Engineer, Tinker Platform at Thinking Machines Lab | CA, US | Rezi

Software Engineer, Tinker Platform at Thinking Machines Lab

Software Engineer, Tinker Platform

Thinking Machines Lab · CA, US

2 days ago

Software Engineer, Tinker Platform

Thinking Machines Lab · CA, US

3 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are hiring a software engineer to build and scale the distributed systems powering Tinker, our post-training platform. You will work on the infrastructure for scheduling jobs across large GPU clusters, maintaining state consistency under failure, and providing a reliable path from idea to a customized model. This role is central to Tinker's reliability and scale, involving system design, operation, and incident response with significant ownership and minimal oversight.

Responsibilities

  • Design, build, and operate the distributed systems underlying Tinker, including job scheduling, resource orchestration, checkpointing, and fault recovery across large multi-node GPU clusters.
  • Own the reliability, latency, and throughput of live production systems serving concurrent training and sampling workloads for external users.
  • Build for graceful degradation and fast recovery: handle node failures, preemptions, and network partitions without losing user progress or data.
  • Improve observability across the platform — metrics, tracing, and alerting that make it possible to detect and diagnose issues before or as they affect users.
  • Lead incident response for the systems you own, drive root-cause analysis, and turn findings into lasting fixes.
  • Partner with research and ML infrastructure teams to understand new training and sampling workloads and to evolve the platform's architecture to support them at scale.
  • Make build-vs-buy and architectural tradeoff calls for core infrastructure components, and mentor other engineers on distributed systems practices.

Requirements

  • 7+ years of experience building, running, and scaling distributed systems in production, with direct on-call ownership of live services.
  • Deep understanding of distributed systems fundamentals: consensus, consistency models, failure detection, scheduling, and state management under partial failure.
  • Track record of designing systems that operate reliably at scale, including diagnosing and resolving production incidents under time pressure.
  • Strong systems-level coding ability in a language such as Python, Go, Rust, or C++, and comfort working close to the infrastructure layer.
  • Experience operating large-scale GPU or accelerator clusters, including job orchestration, resource scheduling, or ML training infrastructure.
  • Familiarity with the training or inference stack for large language models, and an understanding of the systems demands unique to fine-tuning or post-training workloads.
  • Experience with checkpointing, distributed storage, or high-throughput networking (e.g., RDMA, NCCL) in performance-critical settings.
  • Experience building developer-facing infrastructure or APIs where reliability and usability are both first-class concerns.
  • A track record of driving platform-wide architectural decisions and comfort operating with significant autonomy in a fast-changing environment.

Skills

  • Python
  • Go
  • Rust
  • C++
  • Distributed Systems
  • Job Scheduling
  • Resource Orchestration
  • Checkpointing
  • Fault Recovery
  • Reliability
  • Latency
  • Throughput
  • Observability
  • Incident Response
  • Root-Cause Analysis
  • ML Infrastructure
  • GPU Clusters
  • Large Language Models
  • Fine-tuning
  • Distributed Storage
  • High-throughput Networking
  • RDMA
  • NCCL

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • 7+ years

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.
  • We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication.
  • We believe the future worth building is human, and we're hiring people who want to build it.
  • Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs — opening access to capabilities that have previously been concentrated in a handful of labs.
  • We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs.
  • Tinker is rapidly adding new customers, features, and novel use-cases, and the systems underneath it run as a live, always-on production platform.