About the Role
Judgment is the learning infrastructure for AI agents. Agents in production improve from experience, and this role involves building the product experiences and agents that facilitate this loop. You will own problems end-to-end, from customer interaction to product iteration.
Responsibilities
- Shape how the Judgment Agent runs large-scale investigations across thousands of production traces.
- Build the platform for verifying agent changes, including simulated environments and monitors.
- Design interfaces for engineers to understand agent behavior and debugging.
- Design user experiences for engineers to monitor and redirect agent swarms.
- Build workflows to transform production trajectories into datasets, judges, and regression checks.
- Develop platform features such as workspaces, roles, permissions, billing, and usage limits.
- Create an SDK and terminal-first experience for integrating Judgment as a subagent.
Requirements
- Experience building and scaling end-to-end production systems, from data layer to UI.
- Strong technical problem-solving skills, especially in fast-changing, ambiguous environments.
- A builder and tinkerer's mindset with high agency.
- Hands-on experience building with LLMs or agents, or the drive to get there fast.
- Comfort working directly with customers to understand their needs and solve real-world problems.
- Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences.
Location
- San Francisco
Work Type
- On Site
- Full Time
About the Company
- Judgment is the learning infrastructure for AI agents.
