About the Role
You will own AI systems end-to-end, including models, prompts, retrieval, data, evaluations, and debugging. This role involves working across various AI applications such as an ambient scribe, a billing code recommender, medical search, and internal/external agents. Evaluation is a critical component, ensuring the trustworthiness of AI outputs.
Responsibilities
- Build evaluation pipelines to determine AI readiness for deployment, including automated judges, regression suites, human review, and underlying datasets.
- Diagnose and fix AI failure modes by adjusting prompts, retrieval, routing, models, or audio capture.
- Develop production agents, including tool use, orchestration, guardrails, and recovery mechanisms, along with supporting infrastructure like search and vector storage.
- Debug individual visits end-to-end, identify similar cases, quantify the problem, and verify fixes.
- Manage model selection for different requests, implementing weighted routing, staged rollouts, and attribution for safe changes.
- Own AI systems in production, responding to alerts, identifying causes of quality degradation, and deciding on fixes.
- Translate vague clinical complaints into actionable problem statements, metrics, and plans.
- Establish Tali's standards for applied AI and evaluation practices.
Requirements
- 5+ years in production ML, applied AI, or research engineering with experience owning systems for real users.
- Deep experience in evaluation, including building graders, regression suites, or judge pipelines.
- Experience with agentic systems involving multiple models, tool calls, retrieval, and failure recovery.
- Strong systems engineering skills, including backend services, data pipelines, and observability.
- Data-centric approach to improving AI systems through data and feedback loops.
- Proficiency in Python and modern ML tooling.
- Ability to provide and receive candid feedback.
- Skill in making engineering and business cases for complex solutions and owning the results.
- Ability to elevate team members through constructive reviews.
- Experience with speech recognition or real-time audio is a plus.
- Experience in a regulated domain like healthcare or finance is a plus.
- Clinical experience of any kind is a plus.
Skills
- Python
- TypeScript
- GCP
- Cloud Run
- Vertex AI
- LLM
- ASR
Location
- Canada
- US
Work Type
- Full-time
Experience Level
- Senior
- Staff
Salary/Compensations
- $170,000 - $230,000
Benefits
- Flexible work hours
- Comprehensive health and wellness coverage from day one
- Unmetered wellness days
- Competitive PTO
- Winter shutdown (Dec 25 - Jan 1)
- Paid birthdays and Taliversaries
- Extra long long weekends
- $2000 annually in Knowledge Dollars
- Quarterly socials & company outings
About the Company
- Tali AI is a fast-growing startup focused on making healthcare more accessible with AI.
- We are building the clinical operating system, integrating an ambient AI scribe, billing agent, scheduling, and clinical decision support.
- Our platform allows clinicians to focus on patients rather than administrative tasks.
- Thousands of clinicians across Canada and the US use Tali, integrated with various North American health-record systems.
- We have multiple commercial product lines, have documented over 15M patient visits, and saved clinicians over 70 years of time in under three years.
- We operate in a market that rewards speed and requires maintaining clinician trust.
- Tali is one of Linkedin's Top Startups of 2025.
- Tali is part of the renowned Digital Supercluster Project.
