Agent Harness Engineer at Axiom | San Francisco | Rezi

Agent Harness Engineer at Axiom

Agent Harness Engineer

Axiom · San Francisco

3 weeks ago

Agent Harness Engineer

Axiom · San Francisco

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Axiom is building a compounding ecosystem to replace animal testing and reshape clinical trials. We use data to advance ML research and collaborate with AI labs to improve frontier models' ability to reason over our data. This creates a compounding loop where deeper customer understanding shapes data generation, better data improves models and infrastructure, stronger models expand capabilities, and these capabilities are deployed into drug discovery workflows. We are currently focused on solving drug-induced liver injury with an integrated data and agentic system used by top pharma companies and biotechs. Axiom aims to build the world's largest human datasets across major organ systems, paired with an agentic harness to predict human drug outcomes better than animals.

Responsibilities

  • Own the harness: the scaffolding, tooling, and infrastructure that turn frontier models into agents that do long-horizon scientific analysis
  • Build the data backbone: pipelines, storage, and systems for runtime context, agent trajectories, eval results, and training data
  • Build sandboxed execution environments with instant spin-up/tear-down, reproducible and deterministic enough for evals and RL
  • Design and run the eval systems: offline suites, test cases on production traces, LLM-as-judge pipelines, regression gates
  • Work closely with domain experts to encode their taste into measurable rubrics, golden sets, and review workflows
  • Build agent tools and adopt new methods, protocols, and patterns from the state of the art
  • Make every agent run observable and replayable: trace model calls, tool calls, and state transitions, and build debugging tooling
  • Engineer the context: memory, compaction, retrieval, and recovery for coherent long-horizon agent runs
  • Own the loop: retries, budget caps, stop conditions, output verification, permissions, and guardrails
  • Support ML research with environments, reward instrumentation, and rollout infra for RL on agentic tasks

Requirements

  • Engineers who've built with LLM APIs and shipped agentic systems: tool use, loops, and debugging experience
  • Built bespoke evaluation, monitoring, and RL env observability tooling
  • Can tackle deep technical challenges and own/ship simple, clean, maintainable code
  • High ownership: owns outcomes end to end and starts moving without a spec
  • Allergic to complexity: reaches for the simplest system that works and keeps it that way
  • Strong software engineer first, with infrastructure, platform, data, or devtools depth and production systems they're proud of
  • Instinctively asks 'how would we know if this is working?' and builds measurement alongside the feature
  • Reads an agent failure trace like a stack trace
  • Cares about reliability because environments that break silently poison evals and training data
  • Has a knack for surfacing important questions about agent improvement
  • Sharp and confident
  • Keeps up with a technical and sophisticated crowd
  • Comfortable getting in over their head and figuring it out
  • Thrives in a discipline with no playbook, as it's being invented
  • Curious about how things work: engineering/tinkering mindset, good at scavenging the state of the art
  • Passion for learning what 'good' looks like from domain experts and turning it into systems

Skills

  • Python
  • Modal
  • DuckDB
  • FastAPI
  • Docker
  • Containerization
  • Terraform
  • LLM APIs
  • Agentic systems
  • Tool use
  • Loops
  • Evaluation tooling
  • Monitoring tooling
  • RL env observability tooling
  • SvelteKit
  • Svelte 5
  • React

About the Company

  • Axiom is building a compounding ecosystem to replace animal testing and, over time, reshape how clinical trials are run.
  • It starts with deeply understanding the needs of drug hunters inside large pharma.
  • Those needs shape the world-class datasets we build from scratch.
  • We then use that data to advance our own ML research, while also collaborating with leading AI labs to improve frontier models’ ability to reason over Axiom’s data inside Axiom’s agent harness.
  • This creates a compounding loop: deeper customer understanding shapes the data we generate; better data improves frontier models, Axiom’s fine-tuned models, and our agentic infrastructure; stronger models and tooling expand the capabilities we can offer; and those capabilities are forward deployed into pharma's drug discovery workflows, where scientists use them to solve the highest value drug discovery problems.
  • In turn, this helps us identify the next problems to tackle.
  • Today, we are focused on solving drug-induced liver injury through an integrated data and agentic system already being used by 7 of the top 20 pharma companies and several of the world’s most innovative biotechs.
  • Over time, Axiom will build the world’s largest human datasets across all the major organ systems, paired with an agentic harness that uses this data to predict human drug outcomes dramatically better than animals.