About the Role
We are looking for an engineer with strong backend, data, and AI systems experience to build the evaluation and observability foundation for production-grade LLM agents used in complex audit workflows. This role sits at the intersection of backend engineering, data infrastructure, and AI quality, focusing on building evaluation systems that power multimodal retrieval agents and continuously improve quality metrics.
Responsibilities
- Build online and offline evaluation systems for LLM agents, including pipelines that use golden datasets, ground-truth data, human review workflows, and experiment results.
- Create automated quality gates so changes to prompts, context, models, or agent logic can be tested before reaching production.
- Analyze large volumes of agent traces and executions in columnar and analytical databases such as BigQuery or ClickHouse to identify failure modes, quality regressions, latency issues, reliability gaps, and cost optimization opportunities.
- Build reliable data retention and replay mechanisms for long-term analysis of production agent behavior.
- Manage observability tools for tracing, monitoring, debugging, and experiment management of our audit agents.
- Team up with backend engineers to improve the speed and reliability of our retrieval and reasoning agents.
Requirements
- Strong Python and/or backend engineering experience.
- Solid understanding of LLM and agent system evaluation methods (deterministic checks, ground truth, LLM-as-judge, human review, quality metrics) and their appropriate application.
- Experience deploying and operating systems in the cloud, ideally on GCP.
- Hands-on experience building end-to-end retrieval or ML pipeline evaluation systems.
- Experience using LLM observability or experimentation tools such as Braintrust, MLflow, Langfuse, or Weights & Biases.
- Comfortable working with analytical databases, data warehouses, columnar stores, and high-volume event or trace data.
- Understanding of system design, reliability, observability, monitoring, logging, debugging, and operational trade-offs.
- Senior-level engineering judgment, including architectural decision-making and communicating trade-offs.
- Ability to build extensible systems.
- Comfort with ambiguity and reasoning from first principles.
- Excitement to build infrastructure for production AI systems.
Skills
- Backend Engineering
- Data Infrastructure
- AI Systems
- LLM Evaluation
- Observability
- Python
- Cloud Deployment (GCP)
- Retrieval Systems
- ML Pipelines
- LLM Observability Tools
- Analytical Databases
- Data Warehousing
- Columnar Stores
- System Design
- Reliability
- Monitoring
- Debugging
- Data Pipelines
- ETL/ELT Workflows
- Event-Processing Systems
- Feedback Loops
- LLM Optimization
- Workflow Orchestration (Temporal)
Location
- Berlin
Work Type
- On-site
- Full-time
Experience Level
- Senior-level
Benefits
- High impact & growth opportunities
- Mission-driven culture
- Competitive salary
- Significant equity
- Generous coding tools budget
- Flexible vacation
- Team lunches
- Retreats
- Central Berlin office
About the Company
- Cortea is a Berlin startup transforming audits with AI.
- Our AI-powered software and specialized AI agents remove repetitive work, allowing auditors to focus on judgment.
- We are backed by top-tier VCs with over 15m EUR funding.
- We have a working product and paying customers, and are rapidly scaling.
- We value first-principles thinking, speed, trust, and kindness.
- We build side by side in our Berlin office.
Equal Opportunity
- We’re an equal-opportunity team and encourage women and underrepresented groups to apply.
