Staff Software Engineer, Runtime Systems at CoreWeave Europe | GB | Rezi

Staff Software Engineer, Runtime Systems at CoreWeave Europe

Staff Software Engineer, Runtime Systems

CoreWeave Europe · GB

1 weeks ago

Staff Software Engineer, Runtime Systems

CoreWeave Europe · GB

7 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Staff Software Engineer, Runtime Systems role.

Rezi rewrites your resume against CoreWeave Europe's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Staff Software Engineer, Runtime Systems posting at CoreWeave Europe — free, in seconds.

About the Role

We are seeking a Staff Software Engineer, Runtime Systems to help design and build the software and infrastructure that enables demanding AI, simulation, robotics, and engineering workloads to run reliably at scale. This is a hands-on systems engineering role at the intersection of distributed systems, runtimes, workflow execution, programming language concepts, and large-scale compute infrastructure.

Responsibilities

  • Design and build runtime components for complex AI, simulation, and engineering workloads.
  • Define abstractions for workloads, execution environments, dependencies, state, capabilities, and failure.
  • Design interfaces between higher-level services and systems such as Kubernetes, Argo, and OSMO.
  • Establish clear boundaries around which systems own state, decisions, and side effects.
  • Make architectural decisions balancing simplicity, extensibility, performance, and operational reality.
  • Define durable contracts between workload definitions, control-plane services, and execution backends.
  • Develop typed representations and schemas that allow workloads to be transformed safely across systems.
  • Design compatibility and evolution mechanisms for those contracts.
  • Build conformance and validation mechanisms that make guarantees executable rather than dependent on documentation.
  • Reason deeply about retries, partial failure, idempotency, cancellation, dependencies, and uncertain outcomes.
  • Write production-quality software for critical runtime and control-plane components.
  • Build adapters and integrations for heterogeneous execution environments.
  • Diagnose behaviour across application, orchestration, cluster, and infrastructure boundaries.
  • Improve the reliability, observability, and debuggability of distributed workload execution.
  • Work closely with Go, Kubernetes, and infrastructure engineers to turn architecture into production systems.
  • Develop rigorous ways to understand workload performance across large-scale GPU infrastructure.
  • Design experiments that separate real performance gains from noise, warm-up effects, scheduling behaviour, and stragglers.
  • Build repeatable workload and benchmark environments.
  • Use evidence from real execution to challenge assumptions and guide platform development.
  • Lead ambiguous systems problems where the correct architecture is not yet known.
  • Reduce complex problems into smaller contracts and mechanisms that can actually be implemented.
  • Challenge unnecessary abstraction and simplify designs where complexity has outgrown its value.
  • Influence technical direction across teams without requiring direct authority.
  • Mentor engineers and contribute to technical hiring and engineering standards.

Requirements

  • Significant experience building complex systems software, distributed infrastructure, runtimes, workflow systems, or adjacent technology.
  • Deep expertise in at least one of: Distributed systems, Runtime systems, Workflow or execution engines, Programming languages, compilers, or interpreters, Cluster scheduling and orchestration, High-performance or systems software.
  • Strong software engineering fundamentals and production coding ability.
  • Experience designing APIs, protocols, schemas, or contracts between independently evolving systems.
  • Strong understanding of distributed-system failure modes, state, authority, retries, concurrency, and side effects.
  • Strong technical judgement around when abstraction helps and when it simply moves complexity elsewhere.
  • Comfortable entering unfamiliar technical domains and building depth quickly.
  • Strong communication skills and experience influencing architectural decisions across teams.
  • Systems Thinker: You ask where authority lives, what guarantees actually exist, and what happens when systems fail.
  • Technically Deep: You want to understand how systems really behave, not just how they are supposed to behave.
  • Pragmatic: You value elegant engineering, but care more about whether it works for real workloads.
  • Evidence-Driven: You test assumptions and change direction when the evidence says you should.
  • Strong Technical Leader: You can form a view, challenge senior stakeholders, and bring others with you.
  • Commercially Aware: You understand that technical decisions need to improve customer outcomes, engineering velocity, reliability, or economics.
  • Ownership Mentality: You take responsibility for getting difficult systems into production.
  • This position requires access to export controlled information. Applicant must be a U.S. person (U.S. citizen or national, U.S. lawful permanent resident, refugee, or asylee), eligible to access the export controlled information without a required export authorization, or eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency.

Skills

  • Go
  • Rust
  • C/C++
  • Python
  • Kubernetes
  • Containerised infrastructure
  • Argo
  • OSMO
  • Temporal
  • Ray
  • Kubeflow
  • GPU clusters
  • Large-scale AI infrastructure
  • High-performance computing
  • Simulation
  • Robotics
  • Autonomous systems
  • Physical AI
  • Programming language research
  • Compiler research
  • Performance engineering
  • Cloud infrastructure at scale
  • Platforms designed to be operated by autonomous software or AI agents

Location

  • London, SE1 9EA

Work Type

  • Full-time

Experience Level

  • Staff

Salary/Compensations

  • 116,000 GBP to 155,000 GBP

Benefits

  • Family-level Medical Insurance
  • Family-level Dental Insurance
  • Generous Pension Contribution
  • Life Assurance at 4x Salary
  • Critical Illness Cover
  • Employee Assistance Programme
  • Tuition Reimbursement
  • Work culture focused on innovative disruption
  • Discretionary bonus
  • Equity awards

About the Company

  • The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. We are a Living Wage accredited Employer. We work hard, have fun, and move fast. We’re in an exciting stage of hyper-growth, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core, Act Like an Owner, Empower Employees, Deliver Best-in-Class Client Experiences, Achieve More Together. We support entrepreneurial thinking, independent judgement, and collaboration. You’ll work alongside some of the best talent in the industry on technically difficult problems at the frontier of AI infrastructure.

Equal Opportunity

  • CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.