About the Role
We are seeking a network engineer to manage the foundational network layers critical for our large-scale AI training and inference operations. You will ensure the reliability of interconnects across extensive GPU fabrics, including RDMA/RoCE and NVLink/NVSwitch domains. This hands-on role involves debugging production systems, developing essential tooling, and collaborating with cloud providers to resolve network issues, ultimately enabling researchers to trust the underlying infrastructure.
Responsibilities
- Reason about and validate GPU network fabric design across deployments.
- Debug RDMA / RoCEv2 across different NIC vendors.
- Diagnose collective failures of production NCCL, PFC/ECN tuning, and congestion control behavior.
- Own NVLink / NVSwitch interconnect, including fabric manager and IMEX health, link and lane errors, and their interaction with collectives.
- Build host-level network instrumentation and use Linux tooling to create dashboards and alerts.
- Navigate cross-cloud fabric quirks across providers and triage across NIC, driver, kernel, switch, and workload boundaries.
- Drive escalations with cloud-provider networking teams, owning issues end-to-end until resolution.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
- Proficiency in at least one backend language (Python or Rust).
- Experience operating large-scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
- Comfort operating across the stack and owning projects end-to-end.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
- A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
- Fluency with host-level debugging tools on Linux.
- Strong communication skills, internally and with cloud providers.
- Extensive experience with at least one of the following:
- Familiarity with cloud network primitives across at least two cloud providers.
- Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
- Statistical rigor in reliability reasoning — comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
- A track record of writing tooling that made the next debugging session meaningfully faster.
- Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.
- Understanding of deep learning frameworks and their underlying system architectures.
Skills
- Python
- Rust
- Kubernetes
- Slurm
- Linux debugging tools
- NVLink
- NVSwitch
- Fabric manager
- IMEX
- CUDA
- NCCL
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree or equivalent experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Visa sponsorship
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
