About the Role
We are seeking a Senior Machine Learning Engineer to join a high-ownership team focused on delivering production-ready model releases as OEM engagements and release cadence accelerate. This is an applied, delivery-focused role for engineers who enjoy building and iterating on real systems quickly. You will be responsible for taking models from training to product readiness, collaborating with downstream teams to ensure models are deployable on-vehicle and optimizing them for tight runtime constraints using techniques like quantization and distillation.
Responsibilities
- Own end-to-end delivery of model releases, from initial requirements through training, evaluation, iteration, and final readiness for deployment.
- Train and iterate on PyTorch models with a strong experimental approach (hypothesis-driven iteration, ablations, clear evaluation criteria).
- Debug and improve model performance using strong analytical skills—identifying regressions, root-causing issues, and proposing fixes.
- Apply optimisation techniques (e.g., quantisation and distillation where beneficial), understanding trade-offs and when methods are appropriate.
- Collaborate cross-functionally with adjacent ML and performance engineering teams to hand off models, define bottlenecks, and align on optimisation priorities.
- Communicate clearly with stakeholders to align on delivery timelines, trade-offs, and readiness criteria.
Requirements
- Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).
- Strong hands-on experience training and iterating on deep learning models in PyTorch (not just using high-level tooling).
- Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly.
- Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution.
- Familiarity with model optimisation concepts such as quantisation and/or distillation (hands-on is a strong signal, but not a strict requirement if the fundamentals are solid).
- Ability to reason across multiple levels of abstraction—from high-level model behaviour down to practical runtime/latency implications.
- Strong engineering fundamentals and collaboration skills.
- Experience working on models that must meet tight latency / efficiency constraints (edge, embedded, real-time, or similarly constrained production settings).
- Exposure to ML systems spanning training → evaluation → deployment handoff (even if you’re not writing kernels day-to-day).
- Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints.
Skills
- PyTorch
- TensorRT
- CUDA
- Qualcomm QNN
- Triton
- OpenCL
- Quantisation
- Distillation
Experience Level
- Senior
