Machine Learning Engineer, Performance Tooling at Wayve | GB | Rezi

Machine Learning Engineer, Performance Tooling at Wayve

Machine Learning Engineer, Performance Tooling

Wayve · GB

6 days ago

Machine Learning Engineer, Performance Tooling

Wayve · GB

6 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Machine Learning Engineer, Performance Tooling role.

Rezi rewrites your resume against Wayve's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Machine Learning Engineer, Performance Tooling posting at Wayve — free, in seconds.

About the Role

You’ll join the AI Performance Tooling team, making model training and inference faster, more efficient, and more predictable across cloud and embedded hardware. Your mission is to enable data-driven AI performance decisions by building tools that reason across the AI stack, turning profiling data into a clear picture of performance bottlenecks and the impact of proposed changes. You will collaborate with various engineering teams, providing a cross-stack view to translate measurement into actionable recommendations.

Responsibilities

  • Design and build reliable, self-service performance tools that scale across models, hardware targets, and development workflows.
  • Shape how Wayve measures and predicts AI performance, setting standards for other teams.
  • Model theoretical peak performance for a platform, compare it with achieved performance, and pinpoint efficiency losses at the layer and op level.
  • Predict latency, memory, utilization, and compute cost of model or recipe changes before compute is spent.
  • Own monitoring and regression alerting across model builds and training runs.
  • Collaborate with training and runtime engineers to set performance targets and present data-driven cases.

Requirements

  • Deep, hands-on performance engineering in complex systems, including profiling, roofline analysis, latency and throughput optimization, and root-causing performance limitations.
  • A track record of owning a tool or service end-to-end, from design and delivery to adoption by other teams.
  • Strong Python skills and comfort with profiling and instrumenting large production codebases.
  • Hands-on experience developing deep learning models with PyTorch.
  • Data analysis skills to translate noisy measurements into defensible conclusions.
  • Judgment to transform ambiguous performance questions into measurable ones and prioritize effectively.
  • Quantitative communication skills sufficient to influence other teams' priorities.

Skills

  • Python
  • PyTorch
  • Performance Engineering
  • Profiling
  • Roofline Analysis
  • Latency Optimization
  • Throughput Optimization
  • Root-causing Performance Limitations
  • Data Analysis
  • Quantitative Communication

Location

  • In-office

Work Type

  • Full-time
  • Hybrid