About the Role
As a Senior Machine Learning Engineer, you will build and improve the AI systems that power our clinical products. You’ll own complex projects end-to-end, from diagnosing production failures and designing evaluations to building, deploying, and iterating on model and agentic systems. This is a highly hands-on role with significant technical ownership, working closely with clinicians, product managers, and fellow engineers to translate cutting-edge research into reliable, production-grade AI systems.
Responsibilities
- Design and own evaluation pipelines for LLM and agentic systems, combining automated graders, regression testing, production feedback, and human evaluation to measure real product quality.
- Diagnose high-impact failure modes and test improvements across prompting, retrieval, context, routing, data, fine-tuning, or other model and system interventions.
- Develop production systems involving tool use, retrieval, context and state management, routing, orchestration, tracing, and failure recovery.
- Turn production failures and user feedback into better datasets, evaluations, and model behavior through active learning and systematic iteration.
- Distill insights from recent research in LLMs, agents, NLP, speech, and multimodal AI and translate promising ideas into practical experiments.
- Work across models, data, evaluation, orchestration, serving, and observability, while remaining deeply hands-on in code and production debugging.
Requirements
- 5+ years in production ML, research engineering, or applied AI.
- Have built a consequential production AI system or materially improved model behavior in production.
- Strong understanding of modern LLMs, transformers, and production AI systems.
- Experienced designing evaluations for LLMs, agents, or other complex AI systems.
- Can turn ambiguous quality problems into measurable dimensions, datasets, and experiments.
- Familiar with challenges such as grader bias, leakage, misleading aggregate metrics, regression detection, and offline-online mismatch.
- Experience building production systems involving multiple models, tools, retrieval, context, state, routing, or orchestration.
- Understands reliability and failure modes in complex AI workflows, not just individual model calls.
- Proficient in Python and modern ML frameworks; PyTorch preferred.
- Comfortable with deployment, observability, CI/CD, and containerized systems.
- Still highly hands-on: writes code, inspects traces, analyzes failures, and debugs production systems.
- Skilled at building high-quality datasets and feedback loops.
- Experienced using production failures, user feedback, and active learning to improve model and system quality.
- Able to work closely with clinicians, product managers, and fellow engineers.
- Strong communicator who can simplify complex AI concepts for diverse audiences.
- Comfortable owning ambiguous technical problems and driving them to measurable outcomes.
Skills
- LLMs
- Transformers
- Production AI systems
- Evaluation for LLMs and agents
- Python
- PyTorch
- Deployment
- Observability
- CI/CD
- Containerized systems
- Data-centric AI
- Active learning
- Realtime voice
- Conversational AI
- Multimodal systems
- Fine-tuning
- Model adaptation
- Healthcare AI
- Interdisciplinary collaboration
- Mentoring ML engineers
- Open-source contributions
Location
- San Francisco
Work Type
- Hybrid
Experience Level
- 5+ years in production ML, research engineering, or applied AI.
Salary/Compensations
- $225,000 - $300,000, with the addition of significant equity.
Benefits
- Comprehensive medical, dental, and vision coverage for you and your dependents
- 401(k) with a company match of up to 3% of base salary
- Parental leave
- Annual company-wide off-sites, team off-sites and regular team lunches and all-hands gatherings, with travel, lodging and meals covered
- Flexible time off with no annual cap
- Company-wide holidays
- Annual holiday shutdown from December 24–January 1
About the Company
- Ambience is building the AI intelligence platform that restores humanity to healthcare and drives meaningful ROI for health systems across the country.
- Our technology helps providers focus on delivering great care by removing the administrative burden that pulls them away from patients and away from their most impactful work.
- Ambience delivers real-time coding-aware documentation and clinical workflow support across ambulatory, emergency and inpatient settings at the top health systems in North America.
- Our teams operate relentlessly with extreme ownership to build the best solutions for our health system partners.
- We value candor, positivity and deep thought — and we expect a lot from each other because we know the problems we’re solving truly matter.
- Ambience was ranked #1 for Improving the Clinician Experience in the KLAS Research Emerging Solutions Top 20 Report, recognized by Fast Company as one of the Next Big Things in Tech, named one of the best AI companies in healthcare by Inc., and selected as a LinkedIn Top Startup in 2024 and 2025.
- We’re backed by Oak HC/FT, Andreessen Horowitz (a16z), OpenAI Startup Fund, and Kleiner Perkins — and we’re just getting started.
Equal Opportunity
- Ambience Healthcare is an equal opportunity employer and is committed to building a diverse and inclusive workplace.
- We do not discriminate on the basis of race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, genetic information, or any other legally protected status.
- We encourage applicants from all backgrounds to apply.
- Ambience is committed to supporting every candidate’s ability to fully participate in our hiring process. If you need any accommodations during your application or interviews, please reach out to our Recruiting team at accommodations@ambiencehealthcare.com. We’ll handle your request confidentially and work with you to ensure an accessible and equitable experience for all candidates.
