About the Role
Support the creation, implementation, and verification of AI factories, focusing on running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters. Engage with NCCL and collectives to boost performance and scalability, receiving mentorship from senior architects.
Responsibilities
- Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.
- Validate configurations against guidelines for NCCL, collectives, and distributed training frameworks.
- Run key AI/LLM benchmarks — setup, orchestration, result collection, and analysis.
- Investigate and address problems when training jobs or benchmarks fail, hang, or perform below expectations.
- Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.
- Build automation using Python and Shell for conducting benchmarks, retrieving results, and completing regression checks.
- Analyze communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.
- Help identify and recommend improvements to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.
- Work closely with hardware, software, networking, and product teams to prepare AI factories for customer use.
- Contribute to documentation and readiness materials for internal and customer-facing teams.
Requirements
- Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or a related field.
- 5+ years of experience managing Linux-based systems in HPC, distributed systems, or AI/ML environments.
- Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, including some exposure to NCCL.
- Practical knowledge of collective communication patterns like AllReduce and AllToAll, and their application in ML/LLM training.
- Skilled in Python and Shell/Bash for scripting, automation, and tooling.
- Strong communication skills and the ability to work effectively with cross-functional teams.
Skills
- Python
- Shell/Bash
- NCCL
- AllReduce
- AllToAll
- Linux
- AI/ML workloads
- Observability stacks
Location
- Remote
Work Type
- Full-time
Experience Level
- Senior
- 5+ years
Education Level
- Bachelor's degree or equivalent experience
Salary/Compensations
- 152,000 USD - 241,500 USD for Level 3
- 184,000 USD - 287,500 USD for Level 4
Benefits
- Equity
- Benefits
About the Company
- NVIDIA uses AI tools in its recruiting processes.
Equal Opportunity
- NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
