About the Role
This software engineer role will drive AMD’s strategy, architecture, optimization, and tooling to achieve industry-leading AI Pre-training and Distributed Inference Performance on AMD GPUs. You will partner across hardware architecture, AI frameworks, compilers, runtime, ROCm, developer tools, and models to scale performance analysis and optimization.
Responsibilities
- Help with strategy and roadmap for AMD Collectives and Network optimizations.
- Provide guidelines to customers on efficient network load-balancing, workload scheduling, and model sharding strategies.
- Performance tuning, profiling, and analysis of large-scale models for LLM, diffusion, multimodal, RecSys, and generative AI, single node and distributed.
- Explore various tradeoffs and design decisions.
- Participate in hardware-software co-design for future hardware optimizations, especially on scale-up networks, NIC, and scale-out networks.
- Develop and improve framework, tools, and infrastructure for performance estimation, modeling, and reporting.
- Communicate and present the results of the performance analysis and modeling to stakeholders and senior leadership.
- Provide concrete recommendations.
- Cross-team collaboration and working across the organization to identify opportunities and develop strategies.
Requirements
- Deep knowledge with Network, NIC, and GPU hardware architecture, software optimization, performance modeling, AI frameworks, and latest trends in inference and training optimization.
- Hands-on experience in mapping model architecture to low-level software and hardware.
- Understanding the impact of each layer of the stack on model performance.
- Strong knowledge in latest generative model architecture, especially SoTA models.
- Experience with distributed inference and deployment at scale is crucial.
- Multiple years of technical experience in performance optimization.
- Strong technical expertise and experience in performance analysis, projection, and network hardware architecture.
- Deep knowledge and hands-on experience of AI Frameworks such as PyTorch, JAX, vLLM, and SGLang.
- Strong technical leadership skills.
- Ability to work collaboratively with cross-functional teams.
- Ability to coordinate internally and externally.
Skills
- Performance optimization
- Network hardware architecture
- GPU hardware architecture
- Software optimization
- Performance modeling
- AI frameworks
- Inference optimization
- Training optimization
- Generative model architecture
- Distributed inference
- Deployment at scale
- PyTorch
- JAX
- vLLM
- SGLang
- Technical leadership
- Cross-functional collaboration
- Written communication
- Verbal communication
- Presentation skills
Location
- San Jose, CA
Work Type
- Hybrid
Experience Level
- Multiple years of technical experience
Education Level
- PhD or Master's degree in Computer Science, Electrical Engineering, or a related field
Benefits
- AMD benefits at a glance
About the Company
- At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us.
- AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI.
- Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.
Equal Opportunity
- AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.
- We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.
