Reliability Engineer, Supercomputing at Thinking Machines Lab | CA, US | Rezi

Reliability Engineer, Supercomputing at Thinking Machines Lab

Reliability Engineer, Supercomputing

Thinking Machines Lab · CA, US

3 weeks ago

Reliability Engineer, Supercomputing

Thinking Machines Lab · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking an engineer to ensure the reliability of our GPU supercomputing fleet, managing the interface between hardware, firmware, and the operating system. You will identify and resolve long-tail hardware issues, enabling researchers to conduct AI experiments at scale with confidence.

Responsibilities

  • Investigate, reproduce, and remediate issues across large GPU clusters.
  • Own the drivers, kernel surface, and diagnostics spanning hardware, firmware, and OS.
  • Automate the monitoring of fleet reliability and analyze error rates to validate fix effectiveness.
  • Drive the firmware lifecycle, including tracking, qualification, staged rollout, and regression analysis.
  • Engage directly with vendors (GPUs, server OEMs, NICs, storage) to resolve issues and manage RMA processes.
  • Monitor and improve GPU hardware health signals for actionable reliability enhancements.
  • Write clear postmortems and vendor cases to advance issue resolution.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a related field.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large-scale clusters and container orchestration systems (e.g., Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Ability to thrive in a highly collaborative environment with cross-functional partners.
  • A bias for action and initiative to work across different stacks and teams.
  • Fluency with Linux systems and debugging tools.
  • Proven statistical rigor in analyzing reliability.
  • Track record of debugging problems from application symptom to hardware root cause.
  • Comfort reading vendor errata, firmware release notes, and kernel changelogs.
  • Experience engaging hardware vendors directly.
  • Linux kernel literacy, including scheduler, memory management, IRQ paths, and the driver model.
  • Out-of-band management experience (BMC / iDRAC / IPMI / Redfish).
  • Depth in GPU hardware health (Xid error taxonomy, NVLink, NVSwitch, fabric manager, DCGM).
  • Significant ownership of the hardware reliability function at scale.
  • Strong writing skills for vendor cases and postmortems.
  • An instinct for distinguishing between flaky machines, workloads, and tests.

Skills

  • Python
  • Rust
  • Kubernetes
  • Slurm
  • Linux
  • Debugging
  • Statistical analysis
  • Hardware debugging
  • Vendor management
  • Linux kernel
  • Out-of-band management
  • GPU hardware health
  • Xid error taxonomy
  • NVLink
  • NVSwitch
  • Fabric manager
  • DCGM
  • Writing skills

Location

  • San Francisco, California

Work Type

  • Onsite

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship

About the Company

  • Thinking Machines builds AI that extends human will and judgment.