Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure at NVIDIA | Berlin, Germany | Rezi

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure at NVIDIA

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA · Berlin, Germany

1 weeks ago

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA · Berlin, Germany

13 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team is seeking a Senior HPC Cluster Administrator to lead the design, deployment, and reliability of large-scale GPU compute clusters supporting demanding deep learning and HPC workloads. You will drive architectural decisions and collaborate with various teams to ensure infrastructure stays ahead of workload demands.

Responsibilities

  • Own the full lifecycle of GPU compute clusters across heterogeneous Linux environments.
  • Design and scale storage solutions with a roadmap for capacity and performance growth.
  • Lead automation of infrastructure using IaC tools and CI/CD pipelines.
  • Manage and optimize job scheduling via Slurm.
  • Maintain and improve observability stacks and drive proactive incident resolution.
  • Collaborate with ML engineers and software teams to tune cluster configuration.
  • Evaluate and introduce new technologies to improve performance and reliability.
  • Mentor junior engineers and contribute to team-wide engineering standards.

Requirements

  • BS/MS in CS, EE, CE, or equivalent hands-on experience.
  • 5+ years of experience deploying and administering large-scale HPC or ML training clusters.
  • Deep expertise in Linux systems administration at scale.
  • Strong scripting and automation skills in Python and/or bash.
  • Hands-on experience with Slurm.
  • Proficiency with configuration management and IaC (Ansible required; Terraform a plus).
  • Experience with container technologies.
  • Solid understanding of high-speed networking.
  • Experience with distributed/parallel filesystems and storage architecture.
  • Ability to own problems end-to-end and communicate clearly.

Skills

  • Python
  • bash
  • Ansible
  • Terraform
  • Docker
  • Apptainer/Singularity
  • Kubernetes
  • InfiniBand
  • RoCE
  • RDMA
  • EFA
  • NFS
  • Lustre
  • WekaFS
  • GitLab
  • Slurm
  • Prometheus
  • Grafana
  • DCGM
  • PyTorch
  • JAX
  • Megatron
  • BMC
  • IPMI
  • Redfish

Location

  • Poland

Work Type

  • Full-time

Experience Level

  • Senior
  • 5+ years

Education Level

  • BS/MS in CS, EE, CE, or equivalent hands-on experience

Salary/Compensations

  • Poland: 221,250 PLN - 383,500 PLN for Level 3
  • Poland: 292,500 PLN - 507,000 PLN for Level 4

About the Company

  • Join NVIDIA's team of world-class engineers and be part of groundbreaking work.
  • Committed to encouraging a collaborative and inclusive environment where every team member can thrive and make a significant impact.