Software Engineer, ML Infra at Thinking Machines Lab | CA, US | Rezi

Software Engineer, ML Infra at Thinking Machines Lab

Software Engineer, ML Infra

Thinking Machines Lab · CA, US

Yesterday

Software Engineer, ML Infra

Thinking Machines Lab · CA, US

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Software Engineer to interface between research and infrastructure, making critical ML systems decisions and providing hands-on support during high-stakes operations.

Responsibilities

  • Debug across the full stack (kernel, NCCL, scheduler, application, telemetry) to find root causes.
  • Provide embedded, hands-on support during hero runs and major incidents.
  • Serve as the point of contact for researchers when issues arise.
  • Lead postmortems and develop preventative tooling.
  • Mentor other engineers in cross-stack technical breadth.

Requirements

  • Credible, hands-on competence in 4 or more of the following: Linux kernel, networking, GPUs / CUDA, distributed systems runtimes, storage, compilers / language runtimes, observability internals.
  • Comfort operating without a clearly defined owner.
  • Judgment to know when to dig in versus when to escalate.
  • Track record of being the person other engineers escalate to, across 2+ companies.
  • Shipped meaningful contributions in 3+ distinct technical stacks.
  • Track record of leading major incidents where the root cause was non-obvious.
  • Researcher-facing comfort: ability to communicate effectively with researchers and make sound decisions regarding their workloads.
  • Experience operating at the scale of frontier training or inference clusters.
  • Appetite for owning the hardest, least-defined problems in the stack.

Skills

  • Linux kernel
  • Networking
  • GPUs / CUDA
  • Distributed systems runtimes
  • Storage
  • Compilers / language runtimes
  • Observability internals
  • Debugging
  • Incident management
  • Mentoring
  • Communication

Location

  • San Francisco, CA

Work Type

  • Onsite

Experience Level

  • Mid-level
  • Senior

Salary/Compensations

  • $350,000 – $475,000 USD

Benefits

  • Equity
  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.
  • We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication.
  • We believe the future worth building is human, and we're hiring people who want to build it.