About the Role
Join a high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, focusing on open-weights model platforms for inference and fine-tuning. You will work across model serving, inference engines, fine-tuning pipelines, GPU autoscaling, backend services, and observability. This role is for an engineer passionate about optimizing GPU inference and fine-tuning in a rapidly evolving technical landscape.
Responsibilities
- Build infrastructure to move GenAI ideas from prototype to production, increasing AI business impact velocity.
- Work on the open-weights serving stack, including real-time GPU endpoints, high-throughput batch inference, fine-tuning, LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
- Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning.
- Optimize the cost and latency of GPU inference, transforming batch jobs and reducing inference costs.
- Provide product teams with reliable choices across open-weight and closed-source models, including fallback, observability, and cost controls.
- Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.
- Collaborate with ML engineers, product engineers, data scientists, and platform teams to integrate GenAI capabilities into durable platform primitives.
- Shape the future of the centralized GenAI platform, exploring directions like reinforcement learning, agent optimization, and other advanced techniques.
Requirements
- BSc, MSc, or PhD in Computer Science or equivalent.
- 3+ years of industry experience in software engineering.
- Strong backend engineering fundamentals, particularly in Python and distributed systems.
- Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
- Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
- Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production (serving and/or fine-tuning).
- Ability to navigate ambiguous, fast-moving technical areas and translate customer use cases into reusable platform capabilities.
- Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) throughout the software development lifecycle.
Skills
- Python
- Distributed systems
- Production services
- APIs
- Data pipelines
- ML infrastructure
- Observability
- Debugging
- Reliability
- Incident response
- Performance optimization
- Cost optimization
- LLM inference
- LLM fine-tuning
- Open-weight models
- GPU serving
- Batch inference
- Model fine-tuning
- SFT
- DPO
- LoRA
- AI coding tools
Location
- Deliveroo
- DoorDash
- Wolt
Work Type
- Full-time
Experience Level
- 3+ years of industry experience
Education Level
- BSc, MSc, or PhD in Computer Science or equivalent
About the Company
- Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds shared infrastructure for GenAI-powered products.
- The team's mission is to increase the velocity of business impact from GenAI.
- They run frontier open-weight LLMs and VLMs, optimizing GPU serving, batch inference, and fine-tuning for cost and latency.
- The team owns core platform surfaces including LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
Equal Opportunity
- Deliveroo is committed to fostering an environment where everyone can do their best work and feel they belong.
- Equality of opportunity is believed in, and candidates from all backgrounds are welcomed regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.
- Support is offered for individuals with disabilities or long-term health conditions needing assistance with the application process or requiring reasonable adjustments.
