About the Role
We are seeking a seasoned Technical Program Manager with 5-10 years of experience to sustain, scale, and improve operational and software development processes within Compute Systems Engineering and across broader teams, partners, customers, and vendors. This role requires coordination across hardware/software engineering, infrastructure, cloud operations, customer support, and external partners in a fast-moving, deeply technical environment. You will manage complex dependencies, identify risks, improve execution, and establish scalable processes.
Responsibilities
- Adapting systems stack to support the latest hardware platforms and technologies.
- Developing and maintaining performance improvements and customer-facing capabilities.
- Building and operating reliable deployment automation and validation processes.
- Ensuring systems provide sufficient observability across the entire stack.
- Supporting customers with complex technical issues.
- Participating in major hardware bring-ups, delivering critical platform enablers, and conducting system validations.
- Coordinating complex technical programs across multiple teams.
- Identifying dependencies and risks.
- Establishing clear ownership.
- Driving programs toward measurable outcomes.
- Improving execution.
- Establishing processes that continue to work as infrastructure scales.
Requirements
- Demonstrated ability to coordinate complex technical programs across multiple teams.
- Ability to identify dependencies and risks.
- Ability to establish clear ownership.
- Ability to drive programs toward measurable outcomes.
- Strong written, verbal, and interpersonal communication skills.
- Ability to build and maintain productive cross-functional relationships with engineering teams, business partners, customers, and vendors.
- Relevant technical knowledge and strong curiosity about how systems work.
- Ability to understand technical details and communicate effectively with engineers.
- Strong attention to detail, particularly when managing long-running infrastructure programs with numerous technical dependencies, milestones, stakeholders, and operational requirements.
- An open-minded and creative approach to problem-solving.
- Willingness to simplify, adapt, replace, or fundamentally redesign existing processes.
- A hands-on mindset and willingness to contribute directly when needed.
- Ability to assist teams with program execution, technical coordination, documentation, process development, issue investigation, or operational improvements.
- Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
- If you need accommodations during the application process, please let us know.
Skills
- Cloud computing
- High-performance computing
- Accelerated computing
- Large-scale infrastructure
- Device drivers
- Linux kernel
- Virtualization technologies
- Firmware development
- Lifecycle-management processes
Location
- Amsterdam
Work Type
- Full-time
Experience Level
- Approximately five to ten years of relevant experience
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
About the Company
- Nebius is leading a new era in cloud infrastructure for the global AI economy, building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment.
- Built by engineers, for engineers, Nebius addresses challenges across compute, storage, networking, and applied AI.
- Listed on Nasdaq (NBIS) and headquartered in Amsterdam, Nebius has a global footprint with R&D hubs across Europe, the UK, North America, and Israel.
- The team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software, and AI R&D.
- The Compute Systems Engineering organization owns the full systems stack behind Nebius HPC and GPU Cloud deployments, developing and operating core components of the virtualization platform, including Linux kernel, KVM, QEMU, device drivers, and firmware, alongside high-bandwidth networking, NVIDIA GPUs, and other hardware.
- In less than two years, Nebius has successfully operated large-scale B300 and H200 GPU clusters for Cloud customers, comprising tens of thousands of GPUs across multiple data centers in the US, UK, France, and other European countries, with hundreds of thousands more expected.
- A significant challenge is evolving and scaling the Nebius Compute platform to support NVIDIA’s flagship GB300 systems and upcoming VR200 solutions.
- The organization is responsible for adapting systems to new hardware, developing performance improvements, building deployment automation, ensuring observability, and supporting customers with complex technical issues.
- Nebius offers a fast-moving environment with bold thinking, constant growth, meaningful impact, trust, real ownership, and the opportunity to shape the future of AI.
Equal Opportunity
- Nebius is an equal opportunity employer committed to fostering an inclusive and diverse workplace.
- We provide equal employment opportunities in all aspects of employment, without discrimination based on race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
