About the Role
The AI Evaluation Engineer ensures AI agent solutions are accurate, reliable, safe, and production-ready within regulated banking environments. This role provides objective, evidence-based evaluation of AI agent behaviour against defined business, quality, risk, performance, and regulatory criteria, helping ensure AI solutions deliver consistent, trusted outcomes in production. Working within the Reliability Testing phase of the AIOS delivery lifecycle, the role designs, executes, and leads evaluation activities that validate AI agent behaviour across the full development lifecycle. The role helps operationalize AIOS's principle of reliability first, velocity second through disciplined evaluation and objective production-readiness decisions. Ultimately, the role helps ensure AI agent capabilities earn and maintain customer trust by delivering consistent, reliable, and explainable outcomes in production.
Responsibilities
- Design and execute structured evaluation scenarios that validate AI agent accuracy, reliability, safety, compliance, and business outcomes.
- Validate AI agent behaviour against approved business rules, policies, technical requirements, source knowledge, and expected outcomes.
- Conduct regression evaluation across releases and monitor behavioural drift, performance degradation, and newly introduced failure modes.
- Identify, document, prioritize, and track quality issues, defects, and production risks.
- Support production-readiness decisions through objective evaluation evidence and recommendations.
- Analyze evaluation results to identify root causes, recurring quality trends, and opportunities to improve prompts, workflows, knowledge, integrations, and engineering practices.
- Maintain reusable evaluation scenarios, benchmark datasets, expected outcomes, regression suites, and supporting evidence.
- Contribute to continuous improvement of evaluation methodologies, automation, tooling, and engineering feedback loops.
- Support investigation of production issues and validate corrective actions.
- Partner with Agent Engineering, Industry Consultants, Agent Architects, AI Knowledge & Governance, Product, and Delivery teams to ensure evaluation reflects business requirements and production expectations.
- Participate in production-readiness reviews, release planning, and AI agent optimization activities.
- Communicate evaluation findings, quality risks, and recommendations clearly to technical and business stakeholders.
- Lead evaluation activities across one or more AI agent initiatives or squads.
- Review evaluation approaches, production-readiness recommendations, and quality evidence produced by other Evaluation Engineers.
- Coach and mentor less experienced Evaluation Engineers, supporting technical growth and consistent evaluation practices.
- Coordinate evaluation priorities across multiple concurrent initiatives and support delivery planning.
- Drive improvements in evaluation tooling, automation, benchmark management, CI/CD quality integration, and operational effectiveness.
- Analyze systemic quality trends and lead continuous improvement initiatives across multiple AI agent capabilities.
- Partner with Delivery and Engineering leadership to improve quality outcomes, operational consistency, and AI agent reliability.
Requirements
- Typically 3–10+ years of relevant experience in software quality engineering, AI evaluation, AI quality engineering, machine learning evaluation, software testing, or related disciplines.
- Level and scope of responsibility will be determined based on demonstrated technical capability, evaluation expertise, leadership experience, independence, and ability to influence quality outcomes.
- Experience leading evaluation activities, mentoring technical professionals, or coordinating quality initiatives is advantageous for more senior levels.
- Strong understanding of AI evaluation, large language model behaviour, reasoning quality, hallucination detection, safety, instruction adherence, factual accuracy, and business correctness.
- Experience with structured software testing, regression evaluation, production-readiness assessment, and quality engineering.
- Working knowledge of SDLC, CI/CD, automated evaluation, AI observability, and engineering delivery practices.
- Familiarity with benchmark management, evaluation tooling, quality automation, and AI engineering workflows.
- Understanding of privacy, security, governance, and regulatory considerations relevant to enterprise AI.
- Proficiency with Python, SQL, or similar tools supporting evaluation and analysis.
- Experience evaluating LLMs, RAG systems, AI agents, or agentic AI platforms.
- Experience with AI evaluation platforms such as LangSmith, OpenAI Evals, or comparable tools.
- Experience integrating automated evaluation into CI/CD or MLOps workflows.
- Experience with model observability, behavioural-drift detection, or AI production monitoring.
- Banking, financial services, or other regulated industry experience.
- Experience leading technical teams, quality initiatives, or engineering improvement programs.
Skills
- Python
- SQL
Location
- Toronto, Canada
Work Type
- Hybrid
Experience Level
- 3-10+ years
Education Level
- Degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or related discipline, or equivalent practical experience.
Salary/Compensations
- $80,000 - $180,000
Benefits
- Competitive salaries
- Annual bonus potential
- Generous paid time off
- Paid volunteering days
- Wellness benefits
- Robust opportunities for professional growth and career advancement
About the Company
- Zafin is an AI platform company helping regulated institutions modernize how critical work is designed, governed, and delivered.
- Our technology enables organizations to move faster while maintaining the governance, accountability, and control required in highly regulated environments.
- Our portfolio includes Zafin AIOS, an agent orchestration platform for governed AI work; the Zafin Banking Platform, which helps banks modernize product, pricing, offers, billing, loyalty, and relationship management; and Zafin IO, an integration platform that connects data, systems, and workflows across complex enterprise environments.
- Headquartered in Toronto, Canada, Zafin partners with leading financial institutions across North America, Europe, the Middle East, Africa, and Asia-Pacific.
- As AI transforms the future of financial services, we're building the platforms that help regulated organizations adopt AI responsibly and at scale.
Equal Opportunity
- Zafin welcomes and encourages applications from people with disabilities. Accommodations are available on request for candidates taking part in all aspects of the selection process.
- Zafin is committed to protecting the privacy and security of the personal information collected from all applicants throughout the recruitment process. The methods by which Zafin contains uses, stores, handles, retains, or discloses applicant information can be accessed by reviewing Zafin’s privacy policy at https://zafin.com/privacy-notice/. By submitting a job application, you confirm that you agree to the processing of your personal data by Zafin described in the candidate privacy notice.
