About the Role
Judgment is the learning infrastructure for AI agents. You'll build the product experiences that make the agent improvement loop legible and build the agents that run it. This role involves owning problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great.
Responsibilities
- Shape how the Judgment Agent runs large-scale investigations across thousands of production traces.
- Build the platform for verifying agent changes, including simulated environments and monitors for unintended behavior changes.
- Design how engineers understand agent actions and failures through interfaces for long traces, tool calls, and decisions.
- Design how humans watch agent swarms, redirect investigations, and consume findings.
- Build workflows that turn production trajectories into datasets, judges, and regression checks.
- Develop the underlying platform including workspaces, roles, permissions, billing, usage, and limits.
- Create an SDK and terminal-first experience for agents to summon Judgment as a subagent during development.
Requirements
- Experience building and scaling end-to-end production systems, from data layer to UI.
- Strong technical problem-solving skills, especially in fast-changing, ambiguous environments.
- A builder and tinkerer's mindset with high agency.
- Comfort working directly with customers to understand their needs and solve real-world problems.
- Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences.
Skills
- Building and scaling end-to-end production systems
- Technical problem-solving
- Building with LLMs or agents
- Customer interaction
- Communication
About the Company
- Judgment is the learning infrastructure for AI agents.
