About the Role
We are seeking a data scientist to own the evaluation and quality loop of our AutoQA product, focusing on measuring and improving AI system performance. This role offers a steep learning curve and growth opportunities within a fast-growing SaaS company at the intersection of AI and Customer Experience.
Responsibilities
- Translate customer rubrics into AutoQA instructions.
- Design and execute evaluation experiments on large datasets using LangSmith and BigQuery.
- Curate raw production data into golden datasets and generate synthetic data.
- Build LLM-as-a-judge pipelines to assess system quality.
- Make evaluation results actionable by collaborating with the AI team to implement changes.
- Validate retrieval (RAG/IR), tool calling, and speech pipelines.
- Join customer calls to understand quality definitions and provide product feedback.
Requirements
- Solid understanding of how LLMs work and hands-on experience prompting them.
- Good applied statistics knowledge, including experiment design and classifier evaluation.
- Strong Python skills and fluency with the standard data toolkit (Pandas, NumPy, Jupyter).
- An analytical, evidence-first mindset.
- Product sense and empathy for end users.
- Excellent written and verbal communication skills.
- 0 to 2 years of industry experience.
- A team player comfortable wearing multiple hats in a fast-paced environment.
- Experience with LangSmith or similar LLMOps/evaluation tooling is a plus.
- Experience with Google Cloud Platform, BigQuery, Docker, or Kubernetes is a plus.
- Experience with speech/audio processing or ASR evaluation is a plus.
- Experience fine-tuning or building synthetic datasets for LLMs is a plus.
Skills
- LLM Prompting
- Applied Statistics
- Experiment Design
- Classifier Evaluation
- Python
- Pandas
- NumPy
- Jupyter
- SQL
- Product Sense
- Communication
Location
- Amsterdam
- Hybrid
Work Type
- Hybrid
- Full-time
Experience Level
- Entry Level
- Junior
Benefits
- Flexible hours
- Open holiday policy
- Hybrid flexibility
- Visa sponsorship
- Great gear
- Workations
- Team events
About the Company
- Kaizo builds a performance development and quality platform for customer support teams.
- Our AutoQA product uses LLMs to review support conversations against customer quality rubrics.
- We operate a microservices-based stream processing platform handling over 200 million events per day.
- Our LLMOps stack is built around LangSmith for experimentation, prompt management, and tracing.
