Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Orchestration Workload Engineer - ACE - AI Factory role.
Rezi rewrites your resume against Roche's job description. Free.

Tailor your resume to this Orchestration Workload Engineer - ACE - AI Factory role.
Rezi rewrites your resume against Roche's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Orchestration Workload Engineer - ACE - AI Factory posting at Roche — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Orchestration Workload Engineer - ACE - AI Factory posting at Roche — free, in seconds.
About the Role
As a Workload Orchestration Engineer, you will be an expert in workload orchestration, owning and advancing our scheduler tech stack across High-Performance Computing (HPC) platforms. You will drive efficient scheduling, policy management, and resource optimization of multi-node CPU and GPU environments, bridging traditional scientific computing with modern AI paradigms. You will also act as a coach and mentor, solving unique scheduling and infrastructure challenges that impact Roche’s compute architecture.
Responsibilities
- Serve as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC and AI environments.
- Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies.
- Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms.
- Integrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration.
- Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures.
- Lead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams to establish workload orchestration standards, policies, and architectural patterns.
- Act as a technical mentor and coach for junior and mid-level engineers, driving skill development and continuous learning.
- Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization.
Requirements
- Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline.
- Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization.
- Demonstrated track record of leading complex technical initiatives and mentoring engineering peers.
- Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments.
- Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments.
- Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions.
- Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context.
- Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL).
- Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines.
- Apply broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management to solve orchestration problems.
- Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads.
- Strong leadership presence with a dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams.
- Collaborative team player with demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders.
- Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows.
Skills
- SLURM Workload Manager
- HPC
- AI
- Kubernetes
- Singularity/Apptainer
- Containerization
- MPI
- NCCL
- NVIDIA MIG
- InfiniBand
- RoCE
- Ansible
- Terraform
Location
- Kaiseraugst
Work Type
- Hybrid
Experience Level
- Senior
Education Level
- Bachelor's degree
- Advanced degree
About the Company
- At Roche, we are driven by a healthier future to innovate. With over 100,000 employees globally, we advance science and ensure access to healthcare for current and future generations. Our work impacts millions, with over 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We foster an environment that encourages exploration, creativity, and ambition to deliver life-changing healthcare solutions with global impact.
Equal Opportunity
- Roche is an Equal Opportunity Employer.