About the Role
We are seeking a network engineer to manage the foundational network layers critical for large-scale AI training and inference. You will ensure the reliability of multi-thousand-GPU interconnects, including RDMA/RoCE fabric and NVLink/NVSwitch domains. This role involves hands-on, cross-stack debugging, developing essential tooling, and collaborating with cloud providers to resolve network issues, ultimately enabling researchers to trust the underlying infrastructure.
Responsibilities
- Reason about and validate GPU network fabric design across deployments.
- Debug RDMA / RoCEv2 across different NIC vendors.
- Diagnose collective failures of production NCCL, PFC/ECN tuning, and congestion control behavior.
- Own NVLink / NVSwitch interconnect, including fabric manager and IMEX health, link and lane errors, and GPU fabric interaction with collectives.
- Build host-level network instrumentation and use Linux tooling to create dashboards and alerts.
- Navigate cross-cloud fabric quirks across providers and triage across NIC, driver, kernel, switch, and workload boundaries.
- Drive escalations with cloud-provider networking teams, owning issues end-to-end until resolution.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
- Proficiency in at least one backend language (Python or Rust).
- Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
- Comfort operating across the stack and owning projects end-to-end.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
- A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
- Fluency with host-level debugging tools on Linux.
- Strong communication skills, internally and with cloud providers.
- Extensive experience with at least one of the following: Familiarity with cloud network primitives across at least two cloud providers.
- Hands-on experience with NVLink / NVSwitch, fabric manager, and IMEX.
- Statistical rigor in reliability reasoning — comfort reasoning about failure and error rates, distributions, and base rates, and the judgment to separate signal from noise when characterizing a large fabric.
- A track record of writing tooling that made the next debugging session meaningfully faster.
- Familiarity with CUDA/NCCL and performance profiling for distributed training and inference.
- Understanding of deep learning frameworks and their underlying system architectures.
Skills
- Python
- Rust
- Kubernetes
- Slurm
- Linux debugging tools
- Cloud network primitives
- NVLink
- NVSwitch
- Fabric manager
- IMEX
- CUDA
- NCCL
- Performance profiling
- Deep learning frameworks
- System architectures
Location
- San Francisco, California
Work Type
- Onsite
Experience Level
- Mid-level
- Senior
Education Level
- Bachelor's degree or equivalent experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence.
- We're building a future where everyone has access to the knowledge and tools to make AI work for their unique needs and goals.
- We are scientists, engineers, and builders who’ve created some of the most widely used AI products, including ChatGPT and Character.ai, open-weights models like Mistral, as well as popular open source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything.
Equal Opportunity
- As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
- Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.
