About the Role
You'll work closely with product, research-adjacent teammates, and other engineers to ensure agents are reliable, steerable, and trustworthy for real work. This involves improving model behavior and translating those improvements into measurable gains in task completion, reliability, and time savings for users.
Responsibilities
- Design and iterate on agent behavior across real GTM workflows, such as sourcing a Total Addressable Market (TAM) list.
- Map manual, multi-step workflows and transform them into agent-driven flows that match or exceed human performance.
- Build and run evaluations to measure task completion accuracy and identify regressions and failure modes.
- Analyze production failures systematically to improve robustness beyond specific cases.
- Collaborate with product teams to move agent flows from prototype to general availability, defining 'good' for each.
- Build the core agent harness, including memory systems, tool infrastructure, and retrieval architecture.
- Enhance agent performance through prompting strategies, tool-use design, and context construction.
- Design guardrails and safety checks for predictable agent behavior in production.
- Develop a cross-surface evaluation framework for consistent quality measurement across teams.
- Build feedback loops to leverage usage data and production logs for improved prompts, tools, and evaluation coverage.
- Support teams building custom agent variants for specific use cases.
Requirements
- Experience building or shipping production systems with LLMs or agents, including prompting, tool-use design, agent orchestration, retrieval, structured extraction, or fine-tuning.
- Strong backend fundamentals in APIs, databases, and distributed systems for reliable production infrastructure.
- Experience with model or agent evaluation, including designing evaluations, measuring regressions, or converting qualitative feedback into quantitative signals.
- A systems-and-outcomes mindset focused on user product success.
- Comfort debugging real-world failures and a bias for rapid iteration in an evolving field.
Skills
- Prompting strategies
- Tool-use design
- Context construction
- Agent orchestration
- Retrieval
- Structured extraction
- Fine-tuning
- API development
- Database management
- Distributed systems
- Model evaluation
- Agent evaluation
- Systems thinking
- Debugging
- Iterative development
- Agent frameworks
- Tool-calling systems
- Retrieval architectures
- Vector search
- Hybrid search
- RAG
- Eval/benchmark infrastructure
- Fine-tuning in production
- GTM workflows
- Sales workflows
- Marketing workflows
- Lead sourcing
- Audience building
- React
- TypeScript
- Python
- AWS
- Aurora/Postgres
- ECS/Fargate
- Lambda
- OpenSearch
- Elasticache/Redis
- Terraform
- Datadog
- Growth mindset
- Curiosity
- Open-mindedness
Benefits
- Work for free with world-class coaches specializing in creativity, management, and more.
About the Company
- Clay's mission is to help organizations turn any growth idea into reality, viewing growth as a creative practice.
- The company helps thousands of customers, including Anthropic, Notion, Google, and Ramp, with unique data, signals, and AI research.
- Clay raised a $100M Series C in 2025 backed by investors like Sequoia, CapitalG, and First Round, and achieved over $100M in revenue.
- In 2026, the company announced its second employee tender offer in 9 months at a $5B valuation and launched a community equity round.
- The community includes over 11,000 customers, 150+ integration partners, 125+ agencies, 50+ Clay clubs, and 30,000 Slack members.
- Clay fosters a unique culture with team members pursuing diverse interests outside of work.
- Operating principles include negative maintenance and non-attached action.
- The company has been featured in The NYT, Forbes, and First Round Review.
