Orchestration Workload Engineer - ACE - AI Factory at Roche | AG, CH | Rezi

Orchestration Workload Engineer - ACE - AI Factory at Roche

Orchestration Workload Engineer - ACE - AI Factory

Roche · AG, CH

6 days ago

Orchestration Workload Engineer - ACE - AI Factory

Roche · AG, CH

6 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Orchestration Workload Engineer - ACE - AI Factory role.

Rezi rewrites your resume against Roche's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Orchestration Workload Engineer - ACE - AI Factory posting at Roche — free, in seconds.

About the Role

As a Workload Orchestration Engineer, you will be an expert in workload orchestration, advancing our scheduler tech stack across High-Performance Computing (HPC) platforms. You will drive efficient scheduling, policy management, and resource optimization for multi-node CPU and GPU environments, bridging traditional scientific computing with modern AI paradigms. You will also act as a coach and mentor, solving unique scheduling and infrastructure challenges that impact Roche's compute architecture.

Responsibilities

  • Serve as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC and AI environments.
  • Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies.
  • Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms.
  • Integrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration.
  • Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures.
  • Lead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams.
  • Establish workload orchestration standards, policies, and architectural patterns across Roche compute environments.
  • Act as a technical mentor and coach for junior and mid-level engineers.
  • Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization.
  • Deploy policies uniformly using configuration-as-code.

Requirements

  • Bachelor's or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.
  • Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.
  • Demonstrated track record of leading complex technical initiatives and mentoring engineering peers.
  • Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.
  • Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments.
  • Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions.
  • Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context.
  • Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL).
  • Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines.
  • Apply broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management.
  • Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads.
  • Strong leadership presence with a dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams.
  • Collaborative team player with demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders.
  • Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows.

Skills

  • SLURM Workload Manager
  • HPC
  • AI
  • Kubernetes
  • Containerization (Singularity/Apptainer)
  • Infrastructure-as-Code (Ansible, Terraform)
  • GPU scheduling
  • High-speed interconnects (InfiniBand, RoCE)
  • MPI
  • NCCL
  • Observability
  • Capacity Management

Location

  • On-premises

Work Type

  • Hybrid

Experience Level

  • Senior
  • Expert

Education Level

  • Bachelor's degree
  • Advanced degree

About the Company

  • Hosting and Infrastructure (HI) provides mission-critical on-premises infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise.
  • The Value Streams - Accelerated Compute Engineering (ACE) Team acts as a center of excellence and delivery for High Performance Compute and AI Infrastructure across Roche.
  • This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute—helping infrastructure consumers optimize for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI to achieve rapid time-to-value.

Equal Opportunity

  • Genentech is an equal opportunity employer.
  • It is our policy and practice to employ, promote, and otherwise treat any and all employees and applicants on the basis of merit, qualifications, and competence.
  • The company's policy prohibits unlawful discrimination, including but not limited to, discrimination on the basis of Protected Veteran status, individuals with disabilities status, and consistent with all federal, state, or local laws.
  • If you have a disability and need an accommodation in relation to the online application process, please contact us by completing this form Accommodations for Applicants.