Staff ML Performance Engineer (Compiler) at Wayve | London, GBR | Rezi

Staff ML Performance Engineer (Compiler) at Wayve

Staff ML Performance Engineer (Compiler)

Wayve · London, GBR

Today

Staff ML Performance Engineer (Compiler)

Wayve · London, GBR

7 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a Staff ML Performance Engineer, you will optimize ML inference for edge accelerators and GPUs, focusing on running large transformer-based models efficiently on low-cost, low-power edge devices for Wayve's first driving product. You will help set the technical direction for production systems that run reliably on in-vehicle compute, working hands-on across ML systems, compilers, runtimes, kernels, and embedded deployment.

Responsibilities

  • Identify, implement, and validate optimizations in ML compilers, runtimes, and kernels (e.g., operator fusion, scheduling, quantization-aware performance, custom kernels).
  • Profile and pinpoint bottlenecks across the full inference stack (model graph, compiler/runtime, kernel execution, memory movement) and deliver measurable improvements.
  • Build robust benchmarking and regression testing to ensure performance improvements hold across models, devices, and software releases.
  • Develop and optimize for multiple target platforms (e.g., NVIDIA Orin/Thor, Qualcomm), working with cross-functional teams to deliver performant and maintainable solutions.
  • Collaborate with model developers to influence architecture and training/deployment decisions that affect on-device performance.
  • Contribute to technical roadmaps and tooling and help raise the standard of performance engineering across the team.

Requirements

  • Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).
  • Strong proficiency with at least one relevant stack/toolchain (e.g., TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX) and confidence learning adjacent frameworks quickly.
  • Comfort operating at multiple levels of abstraction — from high-level model behavior down to low-level kernel/runtime execution.
  • Strong software engineering fundamentals (debugging, profiling, testing, and maintainable code).
  • Clear communicator and collaborative teammate; able to align multiple stakeholders on performance trade-offs and priorities.
  • Experience with compute graph scheduling and execution on multiple targets.
  • Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints.
  • Experience with NVIDIA and/or Qualcomm SoCs and performance tooling.
  • Python and C++ proficiency.
  • Experience mentoring others and/or driving technical direction in a small, fast-moving team.

Skills

  • ML inference optimization
  • Edge accelerators
  • GPUs
  • Transformer-based models
  • ML systems
  • Compilers
  • Runtimes
  • Kernels
  • Embedded deployment
  • Operator fusion
  • Scheduling
  • Quantization-aware performance
  • Custom kernels
  • Profiling
  • Benchmarking
  • Regression testing
  • NVIDIA Orin/Thor
  • Qualcomm
  • TensorRT
  • CUDA
  • Qualcomm QNN
  • Triton
  • OpenCL
  • MLIR
  • ONNX
  • Software engineering
  • Debugging
  • Testing
  • Maintainable code
  • Python
  • C++

Experience Level

  • Staff