Applied Machine Learning Scientist at Vector Institute | Toronto | Rezi

Applied Machine Learning Scientist at Vector Institute

Applied Machine Learning Scientist

Vector Institute · Toronto

Yesterday

Applied Machine Learning Scientist

Vector Institute · Toronto

2 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As an Applied Machine Learning Scientist, Agent Evaluation and Harness Engineering, you will lead applied research on evaluation, observability, stress-testing, and systematic improvement of AI agents. The role focuses on assessing agent performance and safety across long-horizon, multi-step tasks, and on building methods and tools to help organizations understand whether those systems are working, why they fail, and how to make them measurably better. A core objective is developing adaptive evaluation approaches tailored to Canadian organizations, moving beyond static public benchmarks towards rigorous, organization-specific test environments of end-to-end agentic systems. Working alongside Vector researchers, research professionals, and external partners, the role balances high-quality applied research with the creation of practical technical systems that improve the reliability, safety, security, and effectiveness of deployed agents.

Responsibilities

  • Research and implement state-of-the-art methods for evaluating agents operating over long horizons, multiple tools, changing environments, and partially observable states.
  • Develop evaluations that assess complete agent trajectories, including planning quality, tool selection, intermediate decisions, state transitions, recovery behaviour, verification, termination decisions, resource consumption, and final outcomes.
  • Develop methods for creating organization-specific evaluations from production traces, human feedback, incidents, near misses, support interactions, domain-expert knowledge, and synthetic scenario generation.
  • Create techniques for converting discovered failures into durable regression evaluations that can be rerun across model, prompt, policy, tool, and harness changes.
  • Partner with Vector researchers, Applied ML Specialists, research professionals, and external collaborators to identify consequential agent use cases and create tools, reference agents, and evaluations required for trustworthy deployment.
  • Develop schemas and infrastructure for capturing structured traces of active agents.
  • Research representations of agent trajectories, such as event streams, causal graphs, tool-call graphs, state-transition graphs, and compact trajectory embeddings.
  • Develop approaches for identifying recurrent failure patterns and attributing outcomes to specific components or decisions within an agent system.
  • Build privacy-preserving and security-conscious methods for collecting and analyzing traces in sensitive organizational environments.
  • Research and build agent harnesses incorporating tools, memory, retrieval, sandboxes, permissions, validators, execution loops, recovery strategies, state management, and human approval mechanisms.
  • Develop automated or semi-automated methods for optimizing agent harnesses based on evaluation results and execution traces.
  • Develop safe mechanisms for agents to propose modifications to their own prompts, tools, policies, memory structures, workflow logic, or evaluation criteria while preserving auditability and human control.
  • Lead or contribute to peer-reviewed publications, technical reports, open-source software, benchmark releases, and reference implementations.
  • Contribute to training programs and technical workshops that help Vector partners and external stakeholders design, evaluate, debug, and govern agent systems.
  • Serve as a Vector expert on emerging methods in agent evaluation and harness engineering and connect external stakeholders with relevant members of the Vector research community.
  • Other related duties as assigned from time to time.

Requirements

  • PhD in computer science, computer engineering, machine learning, or a related discipline, or equivalent demonstrated research or engineering experience.
  • Research expertise in one or more of: evaluation of AI agents or language-model systems; automated red-teaming; program synthesis or automated software improvement; AI safety, security, or robustness; multi-agent systems.
  • Strong ability to design controlled experiments and reason about confounding variables, stochasticity, statistical power, evaluator reliability, and reproducibility.
  • Experience evaluating systems whose behaviour unfolds across multiple steps, tool interactions, or environmental state changes.
  • Strong knowledge of Python and experience building high-quality research software.
  • Experience working with modern language models and tool-using agent architectures.
  • Understanding of the distinction between model evaluation and evaluation of the broader model–harness–environment system.
  • Familiarity with open-source machine-learning and agent frameworks such as PyTorch, JAX, Google ADK, LangGraph, the OpenAI Agents SDK, or comparable systems.
  • Comfortable working at the boundary between open-ended research and production-quality engineering.
  • Able to communicate complex findings clearly to technical researchers, engineering leaders, domain specialists, and senior organizational stakeholders.

Skills

  • Python
  • Machine Learning
  • AI Agents
  • Evaluation
  • Observability
  • Stress-testing
  • Agent Performance
  • Agent Safety
  • Long-horizon tasks
  • Multi-step tasks
  • Reliability
  • Security
  • Effectiveness
  • Research Software
  • Language Models
  • Tool-using agent architectures
  • Model evaluation
  • PyTorch
  • JAX
  • Google ADK
  • LangGraph
  • OpenAI Agents SDK

Location

  • Canada

Work Type

  • Full-time

Experience Level

  • PhD or equivalent experience

Education Level

  • PhD

Salary/Compensations

  • $125,800 - $157,300 per year

Benefits

  • Vacation time
  • Floater days
  • GRRSP
  • Health Spending Account
  • Summer Hours program
  • Flexible work arrangements

About the Company

  • Vector Institute believes AI powers possibility by advancing cutting-edge research and translating it into real-world impact through collaboration with research, industry, and government.
  • Vector is committed to fostering a diverse and inclusive culture that reflects its values.

Equal Opportunity

  • The Vector Institute welcomes applications from all qualified candidates, including those who are Indigenous, 2SLGBTQIA+, racialized persons/visible minorities, women, and people with disabilities.
  • If you require an accommodation at any stage of the recruitment or selection process, please contact hr@vectorinstitute.ai. The Vector Institute team will be happy to work with you to ensure your experience is as inclusive and accessible as possible.