About the Role
We are hiring a Centre of Excellence Senior Engineer to help define, standardise, and continuously improve the way Nscale operates its global data centre estate. This is a senior technical role for someone with deep infrastructure operations experience who enjoys solving complex problems, leading operational improvements, and driving consistency across multiple data centre locations. You'll work closely with Infrastructure Operations, Engineering, Deployment, and Platform teams to establish best practices, improve operational efficiency, lead major infrastructure initiatives, and help shape the future operating model for one of the fastest-growing AI infrastructure companies in the world. This role requires occasional travel to Nscale data centres across multiple countries to support deployments, operational improvements, and critical infrastructure initiatives.
Responsibilities
- Lead complex infrastructure projects, including hardware installations, upgrades, and operational improvement initiatives.
- Drive standardisation across Nscale's global data centre estate by developing repeatable operational processes and best practices.
- Identify opportunities to improve operational efficiency, scalability, and service reliability.
- Support the development of Nscale's Centre of Excellence as the operational authority for infrastructure operations.
- Take ownership of complex infrastructure incidents from diagnosis through to resolution.
- Lead major incident response activities and coordinate technical teams during critical events.
- Conduct root cause analysis and implement long-term corrective actions to improve platform resilience.
- Provide technical leadership and mentorship to operations teams across multiple locations.
- Partner with cross-functional teams on capacity planning and infrastructure optimisation initiatives.
- Support operational readiness for new deployments and infrastructure expansion.
- Help develop scalable operational models that support rapid business growth.
- Own the creation and continuous improvement of operational documentation, standards, and runbooks.
- Define and maintain operational procedures for incident management, planned maintenance, and communications.
- Ensure consistency of operational execution across all Nscale data centres.
- Work closely with Infrastructure Engineering, Deployment, Platform Engineering, Network Engineering, and Operations teams.
- Travel to data centre sites to support deployments, resolve complex technical issues, and improve operational excellence.
- Build strong relationships across the business to drive alignment and influence operational best practices.
Requirements
- 5+ years of experience in a technical infrastructure, data centre, or operations environment.
- Extensive experience working within large-scale data centre environments.
- Proven ability to solve complex, systemic infrastructure issues and deliver long-term operational improvements.
- Strong analytical and structured approach to troubleshooting and problem solving.
- Experience leading technical initiatives across multiple sites or operational teams.
- Advanced technical certifications relevant to infrastructure operations.
- Experience supporting hyperscale, HPC, AI, or NVIDIA GPU infrastructure.
- Hands-on server deployment, maintenance, and troubleshooting experience.
- Strong customer service mindset with experience supporting mission-critical infrastructure.
- Highly organised with exceptional attention to detail.
- Comfortable operating in fast-paced, ambiguous environments.
- Strong communication and stakeholder management skills.
- Self-starter with a proactive, ownership-driven mindset.
- Passion for operational excellence and continuous improvement.
- Able to influence without direct authority and build trusted relationships across teams.
Skills
- Infrastructure Operations
- Data Centre Operations
- AI Infrastructure
- Troubleshooting
- Problem Solving
- Technical Leadership
- Capacity Planning
- Infrastructure Optimisation
- Documentation
- Standards
- Governance
- Cross-Functional Collaboration
- Stakeholder Management
Location
- Global
Work Type
- Remote-first collaboration
- Office-based work option
Experience Level
- Senior
Benefits
- Highly competitive package with reviews every 12 months.
- Dynamic progression plan tailored to your ambitions.
- Human-First Flexibility
- Autonomy to shape your day around life's moments.
About the Company
- Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
- At Nscale, our Infrastructure Operations team plays a critical role in maintaining service availability, driving operational excellence, and ensuring our global data centre estate operates safely, reliably, and efficiently.
- We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
Equal Opportunity
- We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
- If there's anything we can do to accommodate your specific situation, please let us know.
