Network Engineer, Supercomputing at Thinking Machines Lab | CA, US | Rezi

Network Engineer, Supercomputing at Thinking Machines Lab

Network Engineer, Supercomputing

Thinking Machines Lab · CA, US

3 weeks ago

Network Engineer, Supercomputing

Thinking Machines Lab · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a network engineer to manage the foundational network layers critical for our large-scale AI training and inference operations. You will ensure the reliability of interconnects across extensive GPU fabrics, including RDMA/RoCE and NVLink/NVSwitch domains. This hands-on role involves debugging production systems, developing essential tooling, and collaborating with cloud providers to resolve network issues, ultimately enabling researchers to trust the underlying infrastructure.

Responsibilities

  • Reason about and validate GPU network fabric design across deployments.
  • Debug RDMA / RoCEv2 across different NIC vendors.
  • Diagnose collective failures of production NCCL, PFC/ECN tuning, and congestion control behavior.
  • Own NVLink / NVSwitch interconnect, including fabric manager and IMEX health, link and lane errors, and their interaction with collectives.
  • Build host-level network instrumentation and use Linux tooling to create dashboards and alerts.
  • Navigate cross-cloud fabric quirks across providers and triage across NIC, driver, kernel, switch, and workload boundaries.
  • Drive escalations with cloud-provider networking teams, owning issues end-to-end until resolution.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large-scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
  • Fluency with host-level debugging tools on Linux.
  • Strong communication skills, internally and with cloud providers.
  • Extensive experience with at least one of the following:
  • Familiarity with cloud network primitives across at least two cloud providers.
  • Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
  • Statistical rigor in reliability reasoning — comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
  • A track record of writing tooling that made the next debugging session meaningfully faster.
  • Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.
  • Understanding of deep learning frameworks and their underlying system architectures.

Skills

  • Python
  • Rust
  • Kubernetes
  • Slurm
  • Linux debugging tools
  • NVLink
  • NVSwitch
  • Fabric manager
  • IMEX
  • CUDA
  • NCCL

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.