About the Role
This role is focused on the end-to-end lifecycle of production-grade AI, from training and fine-tuning specialized models to architecting high-performance inference pipelines. The ideal candidate views AI as a rigorous engineering discipline, building reliable, low-latency, and globally scalable solutions.
Responsibilities
- Lead the adaptation of Large Language Models (LLMs) for domain-specific tasks using techniques like LoRA, QLoRA, and PEFT to balance performance with resource efficiency.
- Architect and optimize inference pipelines to minimize TTFT (Time to First Token) and maximize throughput, implementing quantization, caching strategies, and efficient batching.
- Build and maintain real-time AI pipelines using WebSockets and SSE, ensuring seamless low-latency delivery for voice (ASR/TTS) and text applications.
- Deploy and orchestrate models within containerized microservice architectures (Docker/Kubernetes), ensuring robust monitoring, security, and scalability.
- Work closely with Business Analysts and internal stakeholders to bridge the gap between commercial requirements and technical implementation.
Requirements
- Writing high-quality, maintainable Python code.
- Ensuring solutions are built for reliability, low latency, and global scale.
Skills
- LoRA
- QLoRA
- PEFT
- WebSockets
- SSE
- ASR/TTS
- Docker
- Kubernetes
Location
- London
Work Type
- Full-time
Experience Level
- Mid-level
About the Company
- Join our Global Analytics team.
