About the Role
NVIDIA is seeking outstanding AI Solutions Architects to assist and support customers building solutions with our newest AI and accelerated computing technologies. You will serve as a technical advisor for accelerated systems architecture, GPU and networking systems, cluster design, architectures, orchestration, validation, and production deployment for AI data centers.
Responsibilities
- Partner with ISVs on discovery, architecture reviews, technical deep dives, POCs, benchmarks, demos, and production deployment guidance
- Advise on the design, build-out, and optimization of accelerated AI infrastructure, including large-scale clusters
- Support infrastructure design across compute, networking, storage, containers, observability, security, power, and data center operations
- Drive adoption of systems monitoring, telemetry, and management tools to improve cluster utilization, reliability, performance and workload insight
- Build repeatable reference architectures, deployment guides, sizing guidance, benchmark reports, technical playbooks, demos and whitepapers
Requirements
- BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
- 8+ years of hands-on experience in AI infrastructure, accelerated computing, distributed systems, cloud infrastructure, high-performance computing, or machine learning platforms
- Strong experience designing, deploying, and operating accelerated computing infrastructure at scale
- In-depth knowledge of AI cluster orchestration, scheduling, automation and CI/CD deployment pipelines
- Understanding of data center networking technologies such as InfiniBand, Ethernet, RDMA, network configuration or performance tuning
- Familiarity with infrastructure requirements for AI workloads, including distributed training, inference serving, model deployment, storage performance, and cluster reliability
- Excellent presentation, communication, problem-solving, documentation, and collaboration skills
Skills
- AI infrastructure
- Accelerated computing
- Distributed systems
- Cloud infrastructure
- High-performance computing
- Machine learning platforms
- AI cluster orchestration
- Scheduling
- Automation
- CI/CD deployment pipelines
- InfiniBand
- Ethernet
- RDMA
- Network configuration
- Performance tuning
- Distributed training
- Inference serving
- Model deployment
- Storage performance
- Cluster reliability
- AI factories
- Large GPU clusters
- Multi-node training environments
- Production inference platforms
- LLM training
- Fine-tuning
- RAG
- Inference workflows
- MLPerf
- HPL
- OpenMPI
- NCCL
- GPU communication patterns
- Technical training
- Workshops
- Whitepapers
- Blogs
Experience Level
- 8+ years of hands-on experience
Education Level
- BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
Salary/Compensations
- 184,000 USD - 287,500 USD
Benefits
- Equity
- Benefits
About the Company
- NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.
Equal Opportunity
- NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
