About the Role
As a Software Engineer on Agent Oversight, you will build the platform infrastructure that lets our production agents be observed, evaluated, and improved at scale. This includes building observability tooling, evaluation harnesses, and the pipelines that connect them to improvement loops. You will own your systems end-to-end while maintaining rigorous technical standards.
Responsibilities
- Design and build core platform capabilities for deploying, monitoring, and evaluating agentic applications in production
- Build reliable APIs and data pipelines that capture agent telemetry, evaluation signals, and performance metrics at scale
- Work alongside ML engineers where platform work intersects with evaluation or improvement systems
- Own the reliability, scalability, and observability of platform components serving multiple concurrent enterprise and government customers
- Work cross-functionally with product, forward deployed engineering, and customers to translate real-world deployment requirements into platform features
- Build features end-to-end: system design, implementation, debugging, and testing
- Participate in high-velocity experimentation to validate platform capabilities against real customer usage
Requirements
- 4+ years of professional software engineering experience, with strong fundamentals in backend/distributed systems, APIs, and data pipeline design
- Hands-on experience building production software for ML/LLM-powered products or platforms, such as evaluation systems, observability/monitoring, experimentation infrastructure, agent runtimes, model-serving-adjacent services, or telemetry/data pipelines
- Working knowledge of how LLM or ML systems behave in production: evaluation signals, failure modes, prompt/tool-calling workflows, experiment results, data quality issues, and the tradeoffs between offline evals and live customer behavior
- Experience partnering closely with ML engineers or applied researchers to turn prototypes, eval loops, or model-improvement workflows into reliable platform capabilities
- Experience building infrastructure or platforms that other engineering teams build on top of
- Track record of taking ownership of features or components end-to-end — from design through production — within a larger platform or system
- Comfortable operating in an ambiguous, fast-changing domain where tooling and best practices are still being defined
- Strong problem-solving skills and the ability to work independently or as part of a tight-knit, cross-functional team
- Excited to work directly with ML engineers and customer-facing teams, including challenging assumptions in designs and metrics when platform behavior, model behavior, and customer needs intersect
- Gives direct, substantive feedback on designs and code, and takes it the same way — and mentors others as they grow
Skills
- backend/distributed systems
- APIs
- data pipeline design
- ML/LLM-powered products or platforms
- evaluation systems
- observability/monitoring
- experimentation infrastructure
- agent runtimes
- model-serving-adjacent services
- telemetry/data pipelines
- LLM or ML systems behavior in production
- ML fluency
- observability tooling
- evaluation harnesses
- improvement loops
Location
- San Francisco
- New York
- Seattle
Work Type
- full-time
Experience Level
- 4+ years of professional software engineering experience
Salary/Compensations
- $216,000—$270,000 USD
Benefits
- comprehensive health, dental and vision coverage
- retirement benefits
- a learning and development stipend
- generous PTO
- commuter stipend
About the Company
- Scale’s mission is to develop reliable AI systems for the world's most important decisions.
- As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles.
- We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust.
- At Scale, our mission is to develop reliable AI systems for the world's most important decisions.
- Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact.
- We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.
- We are expanding our team to accelerate the development of AI applications.
Equal Opportunity
- We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace.
- We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.
- We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities.
- We comply with the United States Department of Labor's Pay Transparency provision.
