About the Role
Join our AI Research team as a Data Scientist - AI Evaluation & Benchmarking Manager. Shape and evolve model benchmarking and experimentation capabilities that underpin AI delivery across PwC and its clients. Work within a collaborative applied research environment valuing curiosity, technical rigor, and practical problem-solving. Develop frameworks and experimentation workflows to evaluate emerging AI models, driving evidence-based decisions for client engagements. This role requires technical ownership, thrives in fast-moving environments, and is motivated by building scalable, secure AI research infrastructure.
Responsibilities
- Build and evolve AI benchmarking and experimentation platforms for robust and repeatable model evaluation.
- Design and run end-to-end benchmarking workflows, from understanding client use cases to generating business-ready insights.
- Build scalable evaluation frameworks, metrics, and pipelines, combining hands-on engineering with academic literature review.
- Develop, maintain, and improve experimentation infrastructure for robustness and production-grade engineering.
- Produce clear, client-ready insights and support technical demos and deep dive sessions.
Requirements
- Strong hands-on experience in Data Science concepts or LLM experimentation using structured evaluation frameworks.
- Highly proficient in Python, including asynchronous programming, multithreading, and writing maintainable code.
- Experience deploying ML workloads to cloud platforms (Azure, AWS, or GCP).
- Familiarity with CI/CD and containerization (Docker/Podman).
- Applied knowledge of statistics and experimental design.
- Ability to translate findings into actionable recommendations.
- Comfort managing fast-moving workstreams and operating autonomously.
- Demonstrated emerging leadership behaviors: taking initiative, influencing technical direction, clear communication, and supporting others' development.
- Motivated by ownership and excited by the opportunity to shape AI research platforms impacting client engagements.
Skills
- Data Science
- LLM experimentation
- Python
- Asynchronous programming
- Multithreading
- Maintainable code
- Cloud platforms (Azure, AWS, GCP)
- CI/CD
- Containerization (Docker/Podman)
- Statistics
- Experimental design
- AI evaluation
- Benchmarking
Location
- Office
- Home
- Client site
Work Type
- Hybrid
Benefits
- Empowered flexibility
- Private medical cover
- 24/7 access to a qualified virtual GP
- Six volunteering days per year
