Staff HPC Systems Software Engineer at Nscale | USA | Rezi

Staff HPC Systems Software Engineer at Nscale

Staff HPC Systems Software Engineer

Nscale · USA

1 months ago

Staff HPC Systems Software Engineer

Nscale · USA

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Define the technical direction and evolution of a core HPC platform domain at Nscale, shaping how multiple teams build, automate, and run Slurm-based capabilities within Nscale’s wider cloud-native platform. This high-impact staff-level role combines deep hands-on software engineering with strong systems judgment to ensure Nscale’s HPC services are robust, supportable, and maintainable.

Responsibilities

  • Own and evolve the technical direction for a defined HPC systems domain, such as Slurm platform architecture, scheduler integrations, cluster lifecycle, workload environments, or service automation.
  • Make architectural decisions that balance software quality, operational realities, customer needs, and long-term maintainability.
  • Define how proven Slurm implementations should be packaged, automated, and exposed as a service.
  • Resolve ambiguity around ownership, interfaces, lifecycle boundaries, and operating models across teams.
  • Act as the technical escalation point for the most complex issues within the domain.
  • Establish shared patterns and standards for automation, service lifecycle management, observability, reliability, and supportability across the HPC platform.
  • Drive cross-team design for integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling.
  • Create reusable modules, automation, deployment patterns, and reference implementations that increase engineering leverage.
  • Identify and correct avoidable technical divergence, duplicated effort, and fragile operating models.
  • Ensure domain designs reflect the realities of GPU scheduling, HPC networking, performance isolation, and production operations.
  • Lead technically critical initiatives spanning 2–4 teams or a defined HPC platform area.
  • Unblock delivery by clarifying technical direction and reducing ambiguity in complex system design problems.
  • Contribute hands-on where needed to de-risk or accelerate critical work.
  • Influence engineering teams without formal authority through strong judgement, design clarity, and practical solutions.
  • Partner with adjacent cloud-native software engineers so HPC implementations build on shared platform patterns rather than separate ones.

Requirements

  • Extensive experience designing and building production software and automation for HPC systems, especially Slurm-based environments.
  • Strong track record of writing maintainable, testable, and resilient software in Go, Python, or similar languages.
  • Proven ability to define technical direction across a domain spanning multiple teams or services.
  • Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concerns, and operational trade-offs.
  • Strong practical understanding of GPU-backed infrastructure and HPC networking, including InfiniBand, RoCE, RDMA, and performance-sensitive workload characteristics.
  • Experience integrating HPC systems with cloud-native platforms, APIs, or service delivery models.
  • Experience creating engineering leverage through standards, reusable patterns, shared tooling, and architectural clarity.
  • Strong judgement in balancing short-term delivery with long-term platform health and supportability.
  • Strong written and verbal communication skills, with the ability to align multiple teams around a coherent technical direction.

Skills

  • Go
  • Python
  • Slurm
  • Kubernetes
  • InfiniBand
  • RoCE
  • RDMA

Experience Level

  • Staff-level

Salary/Compensations

  • $225,000—$275,000 USD

Benefits

  • Medical
  • Dental
  • Vision
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation
  • Highly competitive US compensation package (base + bonus + equity)
  • Performance reviews every 12 months
  • Dynamic progression plan tailored to your ambitions
  • Human-First Flexibility

About the Company

  • Nscale is the GPU cloud engineered for AI, providing cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers.
  • Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development.
  • Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
  • We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency.
  • As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work.
  • If you join our team, you’ll be contributing to building the technology that powers the future.

Equal Opportunity

  • We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
  • If there’s anything we can do to accommodate your specific situation, please let us know.