Network Engineer, Supercomputing at Thinking Machines Lab | CA, United States | Rezi

Network Engineer, Supercomputing at Thinking Machines Lab

Network Engineer, Supercomputing

Thinking Machines Lab · CA, United States

1 months ago

Network Engineer, Supercomputing

Thinking Machines Lab · CA, United States

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a network engineer to manage the foundational network layers critical for large-scale AI training and inference. You will ensure the reliability of multi-thousand-GPU interconnects, including RDMA/RoCE fabric and NVLink/NVSwitch domains. This role involves hands-on, cross-stack debugging, developing essential tooling, and collaborating with cloud providers to resolve network issues, ultimately enabling researchers to trust the underlying infrastructure.

Responsibilities

  • Reason about and validate GPU network fabric design across deployments.
  • Debug RDMA / RoCEv2 across different NIC vendors.
  • Diagnose collective failures of production NCCL, PFC/ECN tuning, and congestion control behavior.
  • Own NVLink / NVSwitch interconnect, including fabric manager and IMEX health, link and lane errors, and GPU fabric interaction with collectives.
  • Build host-level network instrumentation and use Linux tooling to create dashboards and alerts.
  • Navigate cross-cloud fabric quirks across providers and triage across NIC, driver, kernel, switch, and workload boundaries.
  • Drive escalations with cloud-provider networking teams, owning issues end-to-end until resolution.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (Python or Rust).
  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
  • Fluency with host-level debugging tools on Linux.
  • Strong communication skills, internally and with cloud providers.
  • Extensive experience with at least one of the following: Familiarity with cloud network primitives across at least two cloud providers.
  • Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
  • Statistical rigor in reliability reasoning — comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
  • A track record of writing tooling that made the next debugging session meaningfully faster.
  • Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.
  • Understanding of deep learning frameworks and their underlying system architectures.

Skills

  • Python
  • Rust
  • Kubernetes
  • Slurm
  • Linux debugging tools
  • Cloud network primitives
  • NVLink
  • NVSwitch
  • Fabric manager
  • IMEX
  • CUDA
  • NCCL
  • Performance profiling
  • Deep learning frameworks
  • System architectures

Location

  • San Francisco, California

Work Type

  • Onsite

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree or equivalent experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
  • We're building a future where everyone has access to the knowledge and tools to make AI work for their unique needs and goals.
  • We are scientists, engineers, and builders who’ve created some of the most widely used AI products, including ChatGPT and Character.ai, open-weights models like Mistral, as well as popular open source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything.

Equal Opportunity

  • As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
  • Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.