About the Role
Join a high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, focusing on the open-weights model platform for inference and fine-tuning. This role is ideal for an engineer passionate about optimizing GPU inference and fine-tuning costs and performance in a rapidly evolving field.
Responsibilities
- Build infrastructure to move GenAI ideas from prototype to production, accelerating business impact.
- Work on the open-weights serving stack, including real-time GPU endpoints, high-throughput batch inference, and fine-tuning.
- Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning.
- Optimize GPU inference cost and latency, reducing job times and inference costs.
- Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, and operational excellence.
- Partner with ML engineers, product engineers, data scientists, and platform teams to create durable platform primitives from GenAI capabilities.
- Shape the future of the centralized GenAI platform, enabling next-generation AI-powered products and agents.
Requirements
- BSc, MSc, or PhD in Computer Science or equivalent.
- 3+ years of industry experience in software engineering.
- Strong backend engineering fundamentals, especially in Python and distributed systems.
- Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
- Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
- Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production.
- Ability to work across ambiguous, fast-moving technical areas and turn customer use cases into reusable platform capabilities.
- Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle.
Skills
- Python
- Distributed systems
- Production services
- APIs
- Data pipelines
- ML infrastructure
- Observability
- Debugging
- Reliability
- Incident response
- Performance optimization
- Cost optimization
- LLM inference
- LLM fine-tuning
- Open-weight models
- GPU serving
- Batch inference
- Model fine-tuning
- AI coding tools
Location
- Remote
Work Type
- Full-time
Experience Level
- 3+ years of industry experience
Education Level
- BSc, MSc, or PhD in Computer Science or equivalent
About the Company
- Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production.
- Our mission is to increase the velocity of business impact from GenAI.
- We run frontier open-weight LLMs and VLMs ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs.
- We own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
Equal Opportunity
- At Deliveroo, we know that a great workplace reflects the world around us and that true diversity and inclusion make us stronger, more creative, and better at what we do. We’re committed to fostering an environment where everyone can do their best work and feel they belong.
- We believe in equality of opportunity and welcome candidates from all backgrounds regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.
- If you have a disability or long-term health condition and need support to apply for one of our roles, or require any reasonable adjustments during the recruitment process, you’ll have the opportunity to let us know once you’ve submitted your application. We’ll share details on how to request support so we can ensure you have a fair and equitable experience.
