Reliability Engineer, Supercomputing at Thinking Machines Lab | San Francisco, California, US | Rezi

Reliability Engineer, Supercomputing at Thinking Machines Lab

Reliability Engineer, Supercomputing

Thinking Machines Lab · San Francisco, California, US

1 months ago

Reliability Engineer, Supercomputing

Thinking Machines Lab · San Francisco, California, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking an engineer to ensure the reliability of our GPU supercomputing fleet, managing the interface between hardware, firmware, and the operating system. Your primary responsibility will be to diagnose and resolve long-tail hardware issues, from NICs and HBM to kernel driver edge cases, ensuring our researchers can conduct AI experiments at scale with confidence.

Responsibilities

  • Investigate, reproduce, and remediate issues across large GPU clusters.
  • Own the drivers, kernel surface, and diagnostics spanning hardware, firmware, and OS.
  • Automate the monitoring of fleet reliability and analyze error rates to validate fix effectiveness.
  • Drive the firmware lifecycle, including tracking, qualification, staged rollout, and regression analysis.
  • Engage directly with vendors (GPUs, server OEMs, NICs, storage) to secure effective fixes and manage RMA processes.
  • Monitor and improve GPU hardware health signals to drive actionable reliability improvements.
  • Write clear postmortems and vendor cases to advance issue resolution.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a related field.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large-scale clusters and container orchestration systems (e.g., Kubernetes or Slurm).
  • Comfort operating across the full stack and owning projects end-to-end.
  • Ability to thrive in a highly collaborative environment with diverse cross-functional partners.
  • A bias for action and initiative to work across different stacks and teams.
  • Fluency with Linux systems and debugging tools.
  • Proven statistical rigor in analyzing reliability.
  • Track record of debugging problems from application symptoms to hardware root causes.
  • Comfort reading vendor errata, firmware release notes, and kernel changelogs.
  • Experience engaging hardware vendors directly.
  • Linux kernel literacy, including scheduler, memory management, IRQ paths, and the driver model.
  • Out-of-band management experience (BMC / iDRAC / IPMI / Redfish).
  • Depth in GPU hardware health (Xid error taxonomy, NVLink, NVSwitch, fabric manager, DCGM).
  • Significant ownership of the hardware reliability function at scale.
  • Strong writing skills for vendor cases and postmortems.
  • An instinct for distinguishing between flaky machines, workloads, and tests.

Skills

  • Python
  • Rust
  • Kubernetes
  • Slurm
  • Linux
  • Debugging
  • Statistical analysis
  • Hardware debugging
  • Vendor management
  • Linux kernel
  • BMC
  • iDRAC
  • IPMI
  • Redfish
  • GPU hardware health
  • Xid error taxonomy
  • NVLink
  • NVSwitch
  • Fabric manager
  • DCGM
  • Writing skills

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
  • We're building a future where everyone has access to the knowledge and tools to make AI work for their unique needs and goals.
  • We are scientists, engineers, and builders who’ve created some of the most widely used AI products, including ChatGPT and Character.ai, open-weights models like Mistral, as well as popular open source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything.

Equal Opportunity

  • As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
  • Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.