About the Role
Define the technical direction and evolution of a core HPC platform domain at Nscale, shaping how multiple teams build, automate, and run Slurm-based capabilities within Nscale’s wider cloud-native platform. This high-impact staff-level role combines deep hands-on software engineering with strong systems judgment to ensure Nscale’s HPC services are robust, supportable, and maintainable.
Responsibilities
- Own and evolve the technical direction for a defined HPC systems domain, such as Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments, or service automation.
- Make architectural decisions that balance software quality, operational realities, customer needs, and long-term maintainability.
- Define how proven Slurm implementations should be packaged, automated, and exposed as a service.
- Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operating models across teams.
- Act as the technical escalation point for the most complex issues within the domain.
- Establish shared patterns and standards for automation, service lifecycle management, observability, reliability, and supportability across the HPC platform.
- Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
- Create reusable modules, automation, deployment patterns, and reference implementations that increase engineering leverage.
- Identify and correct avoidable technical divergence, duplicated effort, and fragile operating models.
- Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation, and production operations.
- Lead technically critical initiatives spanning 2–4 teams or a defined HPC platform area.
- Unblock delivery by clarifying technical direction and reducing ambiguity in complex system design problems.
- Contribute hands-on where needed to de-risk or accelerate critical work.
- Influence engineering teams without formal authority through strong judgement, design clarity, and practical solutions.
- Partner with adjacent cloud-native software engineers so HPC implementations build on shared platform patterns rather than separate ones.
Requirements
- Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
- Strong track record of writing maintainable, testable, and resilient software in Go, Python, or similar languages.
- Proven ability to define technical direction across a domain spanning multiple teams or services.
- Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concerns, and operational trade-offs.
- Strong practical understanding of GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workload characteristics.
- Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
- Experience creating engineering leverage through standards, reusable patterns, shared tooling, and architectural clarity.
- Strong judgement in balancing short-term delivery with long-term platform health and supportability.
- Strong written and verbal communication skills, with the ability to align multiple teams around a coherent technical direction.
Skills
- Go
- Python
- Slurm
- Kubernetes
- InfiniBand
- RoCE
- RDMA
Experience Level
- Staff-level
Salary/Compensations
- $225,000—$275,000 USD
Benefits
- Medical
- Dental
- Vision
- Flexible paid time off
- Parental leave
- Retirement plan participation
- Highly competitive US compensation package (base + bonus + equity)
- Performance reviews every 12 months
- Dynamic progression plan tailored to your ambitions
- Human-First Flexibility
About the Company
- Nscale is the GPU cloud engineered for AI, providing cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers.
- Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development.
- Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
- We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency.
- As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work.
- If you join our team, you’ll be contributing to building the technology that powers the future.
Equal Opportunity
- We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
- If there’s anything we can do to accommodate your specific situation, please let us know.
