About the Role
We are seeking a disciplined and dynamic Systems Engineer focused on server CPU-based systems to join our compute rack validation team. This role involves leading server rack and blade hardware systems deployment, installation, and inventory management. You will drive post-silicon validation throughout the program lifecycle, innovate system bring-up and enablement, and ensure the delivery of high-quality technologies. Your technical leadership, systems engineering, and hardware expertise will be crucial for product development, root cause analysis, and resolution. Agility and collaboration are essential for working within System Validation and other engineering teams.
Responsibilities
- Install, configure, commission, and decommission blade servers, chassis, switches, and supporting infrastructure.
- Lead rack and stack activities, including equipment mounting, cable management, and labeling.
- Execute hardware upgrades, replacements, and troubleshooting of server and network components.
- Maintain accurate asset records within DCIM platforms and inventory management systems.
- Conduct physical audits and reconcile inventory discrepancies.
- Track hardware movements, deployments, and decommissions through change management processes.
- Document installation procedures, rack layouts, cabling diagrams, and inventory updates.
- Support data center migration, expansion, and refresh projects.
- Collaborate with engineering, operations, logistics, and project management teams.
- Adhere to all data center safety, security, and operational standards.
- Develop, set up, and scale methodologies for at-scale test execution, lab hardware and system software capabilities, and system visibility/debug tools for AI compute rack bring-up and validation.
- Work independently in a production-ready environment, maintaining accurate inventory and asset records.
- Triage issues during server rack validation, bring-up, Post-Silicon Validation, and production phases, ensuring timely and quality resolution.
- Lead test execution for key domains within AI compute solutions, including CPU, GPU, memory, HBM, and IO.
- Drive technical innovation to improve system validation capabilities, including tools, script development, methodology enhancement, and cross-functional initiatives.
Requirements
- Strong analytical and problem-solving skills with attention to detail.
- Experience in Blade server installation and maintenance (Cisco UCS, HPE Synergy, Dell MX, or similar).
- Rack and stack deployments in enterprise or hyperscale environments.
- Copper and fiber cabling installation and management.
- Experience with DCIM and asset management platforms.
- Strong understanding of server, storage, and networking hardware.
- Experience performing inventory audits and maintaining asset accuracy.
- Ability to read rack elevation diagrams, cabling schematics, and deployment documentation.
- Familiarity with ticketing and change management systems.
- Exposure to Linux (Ubuntu) OS bootable images and system firmware basics for image building, provisioning, and firmware flashing.
- Exposure to automation testing for hardware acceptance and best-known-config testing.
- Exposure to Python script development and execution.
- Proven experience in understanding, defining, and enabling storage and networking capabilities in a lab environment for rack and blade validation.
- Excellent communication and coordination skills.
- Detail-oriented, highly organized, able to prioritize, and manage multiple work streams to deadlines.
- Technical leadership capabilities to champion improvements in platform validation.
- Experience working with data center technical staff, 3rd party vendors, and ODMs throughout server system product development.
- Self-starter, able to independently drive tasks to completion.
Skills
- Server CPU based system focus
- Server rack and blade hardware systems deployment
- Hardware installation
- Inventory management
- Post-silicon validation
- System bring-up
- System enablement
- Silicon validation
- System validation
- Hardware bring-up
- Hardware validation
- Hardware debug
- Root cause analysis
- Problem-solving
- Blade server installation and maintenance
- Rack and stack deployments
- Cable management
- DCIM platforms
- Asset management systems
- Inventory audits
- Change management processes
- Data center migration
- Data center expansion
- Data center refresh projects
- Data center safety standards
- Data center security standards
- Data center operational standards
- At-scale test execution
- System software capabilities
- System visibility
- Debug tools
- AI compute rack validation
- Triage
- CPU validation
- GPU validation
- Memory validation
- HBM validation
- IO validation
- Technical innovation
- Tools development
- Script development
- Methodology enhancement
- Cross-functional initiatives
- Analytical skills
- Attention to details
- Cisco UCS
- HPE Synergy
- Dell MX
- Copper cabling
- Fiber cabling
- Server hardware
- Storage hardware
- Networking hardware
- Rack elevation diagrams
- Cabling schematics
- Deployment documentation
- Ticketing systems
- Linux (Ubuntu)
- System firmware
- Image building
- Provisioning
- Firmware flashing
- Automation testing
- Python script development
- Storage system enablement
- Network system enablement
- Communication skills
- Coordination skills
- Organization
- Prioritization
- Work stream management
- Technical leadership
- Platform validation improvements
- Data center technical staff collaboration
- 3rd party vendor collaboration
- ODM collaboration
- Self-starter
- Masters or PhD in Electrical Engineering, Computer Engineering or related field
- Complex systems engineering challenges
- HW-FW-SW debug
- AI/ML rack scale systems design and deployment
- Industry standards for hardware development
- Best practices for hardware development
- Emerging technologies in AI
- Emerging technologies in Data Center infrastructure
- ODM partner engagement
- Staffing vendor engagement
Location
- Austin, TX
Work Type
- Full-time
Experience Level
- 10+ years of work experience
Education Level
- Masters or PhD in Electrical Engineering, Computer Engineering or a related field
Benefits
- Medical coverage
- Dental coverage
- Vision coverage
- Flexible Spending Accounts (FSAs)
- Health Savings Accounts (HSAs)
- Disability insurance
- Life insurance
- 401(k) retirement plan
- Commuter benefits
- Wellness services
- Employee Assistance Programme (EAP)
About the Company
- Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future.
- We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone.
- We offer an equal opportunity process and understand that there are visible and invisible differences in all of us.
- We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
Equal Opportunity
- We offer an equal opportunity process and understand that there are visible and invisible differences in all of us.
- We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
