Lead Engineer for Manufacturing and Datacenter Lab, Trainium Manufacturing, Quality and Reliability
Amazon · Austin, TX, US
20 hours agoImpress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Lead Engineer for Manufacturing and Datacenter Lab, Trainium Manufacturing, Quality and Reliability role.
Rezi rewrites your resume against Amazon's job description. Free.

Tailor your resume to this Lead Engineer for Manufacturing and Datacenter Lab, Trainium Manufacturing, Quality and Reliability role.
Rezi rewrites your resume against Amazon's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Lead Engineer for Manufacturing and Datacenter Lab, Trainium Manufacturing, Quality and Reliability posting at Amazon — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Lead Engineer for Manufacturing and Datacenter Lab, Trainium Manufacturing, Quality and Reliability posting at Amazon — free, in seconds.
About the Role
We are establishing a critical new function that bridges manufacturing outcomes with datacenter operational performance. This role will build and lead this strategic capability, serving as the essential feedback loop between manufacturing operations and AWS datacenter fleet performance. You will establish and operate a specialized preparedness lab focused on analyzing datacenter performance of manufactured systems to identify root causes of field rework and repairs, feeding critical insights back into manufacturing processes, test strategies, and design improvements.
Responsibilities
- Own operational production performance of Trainium systems across the entire product lifecycle from manufacturing through datacenter deployment and fleet operations.
- Design and build a preparedness lab replicating datacenter conditions for assembly, repair, and system testing.
- Define and drive assembly and repair recipes in the manufacturing lab as the baseline prior to high volume manufacturing and datacenter deployment.
- Ensure all manufacturing and datacenter test flows are regressed in the manufacturing lab prior to deployment.
- Influence hardware design strategy for Design for Manufacturing (DFM), Design for Reliability (DFR), and Design for Test (DFT) based on field failure analysis.
- Establish data-driven analytics frameworks connecting manufacturing test data to datacenter performance, leveraging ML techniques to predict field failures.
- Build and mentor a cross-functional team spanning manufacturing, test, quality, and reliability engineering; perform technical promotion assessments as a force multiplier.
- Collaborate with AWS datacenter operations teams to understand failure modes, repair patterns, and operational challenges firsthand; translate operator insights and field learnings into actionable manufacturing process improvements and design changes.
- Drive continuous improvement reducing failure rates and lifecycle degradation through rapid feedback loops.
- Develop or adapt manufacturing processes at the ODM and CM, including defining fixture requirements, critical assembly requirements, test methodology, signal integrity, power and heat management requirements.
Requirements
- Experience carrying design concepts through exploration, development, and into deployment or mass production.
- Experience communicating with customers, technical, regulatory, business teams, and management to collect requirements, describe product features, and technical designs.
- Experience communicating results to senior leadership, or experience communicating complex information and solutions to senior stakeholders and influencing decisions.
- 8+ years industry experience in one or more of the following: Manufacturing Engineering, Test Engineering, Quality Engineering, Reliability Engineering, or Datacenter Infrastructure Engineering.
- 7+ years working directly with engineering teams in cross-functional environments.
- Experience with AI/ML acceleration systems, high-performance computing servers, or complex multi-rack systems.
- Demonstrated track record delivering stable, performant hardware solutions meeting cost and quality targets.
- Experience with System Mechanical & Thermal design for air-cooled and liquid-cooled systems.
- Strong problem-solving capabilities to isolate, define, and resolve complex problems spanning manufacturing quality and field reliability.
- Experience with root cause analysis methodologies (8D, 5-Why, Fishbone, FMEA) and implementing corrective/preventive actions.
- Proficiency in data analysis tools, statistical methods, and programming (Python, Bash, Shell script, Linux).
- Experience working with ODMs, JDMs, component vendors, and internal design teams on cross-boundary triaging, debugging, and resolving issues.
- Experience in Design for Manufacturing (DFM), also known as Design for Manufacturability, a product design approach that focuses on optimizing the ease and cost of manufacturing a product.
- Can be given complex hardware engineering problems to solve and design project strategy that splits work appropriately for parallel development.
Skills
- AI/ML acceleration systems
- High-performance computing servers
- Complex multi-rack systems
- System Mechanical & Thermal design
- Root cause analysis methodologies (8D, 5-Why, Fishbone, FMEA)
- Data analysis tools
- Statistical methods
- Programming (Python, Bash, Shell script, Linux)
- Design for Manufacturing (DFM)
Location
- Austin, Texas
Work Type
- Onsite
Experience Level
- 8+ years industry experience
- 7+ years working directly with engineering teams
Education Level
- BS or MS degree in Electrical Engineering, Mechanical Engineering, Computer Engineering, Industrial Engineering, or related technical fields
Salary/Compensations
- 159,200.00 - 215,300.00 USD annually
Benefits
- Sign-on payments
- Restricted stock units (RSUs)
- Health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage)
- 401(k) matching
- Paid time off
- Parental leave
About the Company
- Annapurna Labs is a wholly owned subsidiary of AWS, focused on developing custom silicon and servers including the Nitro(K2), Graviton, Inferentia, and Trainium families of processors.
- Machine Learning Annapurna functions as a vertically integrated team including software, firmware, hardware, and silicon design in a single organization.
- We are the Trainium Servers and Systems organization under MLA focused on Hardware Development, Software Development, Fleet Ops Systems, and Manufacturing, Quality, and Reliability.
- This position is in the Manufacturing, Quality and Reliability team.
Equal Opportunity
- Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
- Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.