About the Role
As a Machine Learning Engineer on Agent Oversight, you will drive the end-to-end lifecycle that ensures our production agents perform reliably and improve over time. This includes building observability tools, designing robust evaluation frameworks, and developing improvement loops. You will navigate the entire ML loop while maintaining rigorous technical standards.
Responsibilities
- Build or contribute to observability into agent behavior in production.
- Design evaluation methodologies and metrics for agentic applications, and work with the platform to make them run automatically, at scale.
- Build, ship, and own ML systems that detect drift, anomalies, or misalignment in production agent behavior.
- Design and run rigorous experiments to validate model and agent performance improvements before they ship.
- Work alongside software engineers on the platform where your work intersects with broader infrastructure.
- Collaborate closely with product managers, customers, data annotators, Forward Deployed Engineers, and other engineering teams to translate enterprise and government requirements into robust platform capabilities.
- Contribute to novel methods and approaches that push the state of the art for agent evaluation and improvement, or focus on building ML systems that hold up reliably at scale in production.
Requirements
- 5+ years of experience as an ML engineer or applied scientist, ideally on a production ML or LLM-powered system.
- Strong grounding in at least two of the following: Building or scaling evaluation, monitoring, or continuous-learning infrastructure for ML/agentic systems; Design experience for agent systems (architecture, orchestration, tool use); Developing new methods, reward models, or model training/fine-tuning approaches.
- Hands-on experience with LLMs and agent architectures — tool use, planning, multi-agent orchestration.
- Comfortable partnering with software engineers to productionize research and experimental work.
- Rigorous approach to experimentation: clear hypotheses, real statistical grounding, and results that hold up under scrutiny.
- Track record of collaborating across functions (Product, Forward Deployed Engineering, etc.) to navigate ambiguous requirements and bring them to production.
- Gives direct, substantive feedback on designs and code, and takes it the same way — and mentors others as they grow.
Skills
- LLMs
- agent architectures
- tool use
- planning
- multi-agent orchestration
- ML systems
- evaluation
- monitoring
- continuous-learning infrastructure
- agent systems design
- reward models
- model training
- fine-tuning approaches
- RLHF
- SFT
- reward modeling
- verifiable-reward systems
- model optimization
- systems optimization
- latency optimization
- cost optimization
- inference efficiency
Location
- San Francisco
- New York
- Seattle
Work Type
- full-time
Experience Level
- 5+ years of experience
Salary/Compensations
- $216,000—$270,000 USD
Benefits
- comprehensive health, dental and vision coverage
- retirement benefits
- a learning and development stipend
- generous PTO
- commuter stipend
About the Company
- Scale’s mission is to develop reliable AI systems for the world’s most important decisions.
- As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles.
- We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust.
- At Scale, our mission is to develop reliable AI systems for the world's most important decisions.
- Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact.
- We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.
- We are expanding our team to accelerate the development of AI applications.
Equal Opportunity
- We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace.
- We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.
- We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities.
- We comply with the United States Department of Labor's Pay Transparency provision.
