About the Role
Bridge the gap between cutting-edge AI research and production deployment by designing, evaluating, and deploying intelligent systems that combine frontier models with structured knowledge, retrieval, traditional machine learning, and enterprise software. Work across a diverse portfolio of AI challenges spanning multiple industries and use cases, building systems that create measurable business impact.
Responsibilities
- Design and deploy production AI agents that leverage the latest advances in large language models, reasoning, retrieval, memory, and tool use.
- Architect intelligent systems that combine LLMs, traditional machine learning, structured knowledge, enterprise data, and deterministic software into reliable production workflows.
- Engineer customer intelligence layers, retrieval pipelines, memory systems, and knowledge representations that allow agents to reason over large, heterogeneous enterprise data.
- Develop multi-agent systems that coordinate reasoning, planning, tool execution, and human oversight.
- Translate frontier AI research into production systems by rapidly evaluating new models, prompting techniques, reasoning paradigms, and agent architectures.
- Own the full experimentation lifecycle, from hypothesis generation to production rollout.
- Design rigorous evaluation frameworks using offline benchmarks, online A/B experiments, golden datasets, regression suites, LLM-as-a-Judge, and human evaluation.
- Run controlled experiments and ablation studies to understand the contribution of different models, prompts, retrieval strategies, reasoning techniques, memory systems, and agent architectures.
- Continuously evaluate newly released frontier models and determine where they meaningfully improve quality, latency, reliability, or cost.
- Develop confidence estimation, reflection, and continuous learning systems that improve agents over time using real-world feedback.
- Measure success through business outcomes, not benchmark scores.
- Build production-quality AI systems with a strong emphasis on reliability, observability, latency, safety, and cost.
- Design agent guardrails, fallback strategies, tracing, monitoring, and evaluation pipelines that enable safe deployment in high-stakes environments.
- Collaborate with infrastructure engineers to deploy AI systems securely within enterprise cloud environments.
- Build human-in-the-loop workflows that effectively combine AI automation with expert oversight.
- Partner directly with enterprise customers to understand their business, data, and operational challenges.
- Translate ambiguous customer problems into production AI architectures.
- Rapidly prototype new ideas, validate them with customers, and evolve successful solutions into scalable production systems.
- Identify reusable patterns that become core capabilities across many enterprise deployments.
Requirements
- 5+ years of software engineering, machine learning, or applied AI experience.
- Strong Python programming skills.
- Experience building production AI systems using LLMs.
- Experience with modern AI tooling, including OpenAI, Claude, MCP, agent frameworks, vector databases, or retrieval systems.
- Strong understanding of machine learning fundamentals and modern language models.
- Experience designing or evaluating AI systems using quantitative metrics.
- Excellent communication skills and the ability to work directly with enterprise customers.
- Experience building production AI agents or autonomous systems.
- Deep understanding of reasoning, retrieval, memory, planning, and tool use.
- Experience designing evaluation frameworks for LLMs and agentic systems.
- Experience with RAG, semantic search, knowledge graphs, customer intelligence systems, or structured knowledge representations.
- Experience with fine-tuning, distillation, reinforcement learning, small language models, or model optimization.
- Familiarity with multimodal AI systems and frontier foundation models.
- Experience building distributed production systems.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Experience with Docker, Kubernetes, CI/CD, and production observability.
- Experience integrating AI systems into enterprise software environments.
- Experience working directly with enterprise customers.
- Ability to translate ambiguous business problems into technical architectures.
- Strong written and verbal communication skills.
- Experience leading technical workshops, architecture reviews, or customer design sessions.
Skills
- Python
- LLMs
- OpenAI
- Claude
- MCP
- Agent frameworks
- Vector databases
- Retrieval systems
- Machine learning fundamentals
- Modern language models
- Quantitative metrics
- Communication skills
- Reasoning
- Memory
- Planning
- Tool use
- RAG
- Semantic search
- Knowledge graphs
- Customer intelligence systems
- Structured knowledge representations
- Fine-tuning
- Distillation
- Reinforcement learning
- Small language models
- Model optimization
- Multimodal AI systems
- Frontier foundation models
- Distributed production systems
- AWS
- Azure
- GCP
- Docker
- Kubernetes
- CI/CD
- Production observability
- Enterprise software integration
Location
- San Francisco
- New York
- Seattle
Work Type
- Full-time
Experience Level
- Senior
- 5+ years
Salary/Compensations
- $216,000—$270,000 USD
Benefits
- Base salary
- Equity
- Comprehensive health, dental and vision coverage
- Retirement benefits
- Learning and development stipend
- Generous PTO
- Commuter stipend
About the Company
- Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications.
- We partner with the world's leading enterprises and government organizations to accelerate their AI transformation through frontier AI systems that solve real business problems.
- Every day, we work with organizations across finance, healthcare, manufacturing, media and telecommunications to build production AI agents that automate complex workflows, help humans, reason over enterprise knowledge, and operate safely at scale.
- At Scale, our mission is to develop reliable AI systems for the world's most important decisions.
- Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact.
- We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force.
- We are expanding our team to accelerate the development of AI applications.
Equal Opportunity
- We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace.
- We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.
- We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities.
- We comply with the United States Department of Labor's Pay Transparency provision.
