About the Role
As a Staff ML Performance Engineer, you will optimize ML inference for edge accelerators and GPUs, focusing on running large transformer-based models efficiently on low-cost, low-power edge devices for Wayve's first driving product. You will help set the technical direction for production systems that run reliably on in-vehicle compute, working hands-on across ML systems, compilers, runtimes, kernels, and embedded deployment.
Responsibilities
- Profile and pinpoint bottlenecks across the full inference stack and deliver measurable improvements.
- Implement and validate optimisations in compilers, runtimes, and/or kernels.
- Build robust benchmarking and regression testing to ensure performance improvements hold across models, devices, and software releases.
- Optimise for multiple targets and work with teams to support these in a maintainable way.
- Collaborate with model developers to influence architecture and training/deployment decisions that affect on-device performance.
- Contribute to technical roadmaps and tooling and help raise the standard of performance engineering across the team.
Requirements
- Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).
- Strong proficiency with at least one relevant stack/toolchain and confidence learning adjacent frameworks quickly.
- Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution.
- Strong software engineering fundamentals (debugging, profiling, testing, and maintainable code).
- Clear communicator and collaborative teammate; able to align multiple stakeholders on performance trade-offs and priorities.
- Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints.
- Experience with NVIDIA and/or Qualcomm SoCs and performance tooling.
- Python and C++ proficiency.
- Experience mentoring others and/or driving technical direction in a small, fast-moving team.
Skills
- ML inference
- Edge accelerators
- GPUs
- Transformer-based models
- ML systems
- Compilers
- Runtimes
- Kernels
- Embedded deployment
- TensorRT
- CUDA
- Qualcomm QNN
- Triton
- OpenCL
- Python
- C++
Location
- London
Work Type
- Full-time
- Hybrid
Experience Level
- Staff
About the Company
- Wayve is focused on enabling the first driving product through efficient ML inference on edge devices.
