Software Engineer, AI Evaluation at Nuna | CA, US | Rezi

Software Engineer, AI Evaluation at Nuna

Software Engineer, AI Evaluation

Nuna · CA, US

1 months ago

Software Engineer, AI Evaluation

Nuna · CA, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Software Engineer, AI Evaluation role.

Rezi rewrites your resume against Nuna's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Software Engineer, AI Evaluation posting at Nuna — free, in seconds.

About the Role

Nuna is building an AI health coach to help patients manage chronic conditions at home. This role involves owning the end-to-end evaluation system for AI agents, ensuring they are safe and effective before deployment.

Responsibilities

  • Build testing harnesses and evaluation infrastructure for agentic products and internal tooling.
  • Own the architecture and content of AI agent evaluations end-to-end.
  • Ensure all agentic deployments pass through the testing apparatus before shipping.
  • Own release gates to prevent unsafe or low-quality behavior from reaching patients.
  • Build ground truth, judges, and metrics, validating evaluation trustworthiness.
  • Build functional tooling for labeling and review workflows for non-engineers.
  • Facilitate the loop from evaluation results to model and prompt refinement for safe iteration.

Requirements

  • Significant experience building and shipping reliable production systems and tooling.
  • Deep understanding of AI system evaluation methods (LLM-as-judge, red-teaming, synthetic scenario generation, multi-turn and agentic evaluation).
  • Experience deploying production evals and automated AI tooling.
  • A testing mindset applied to building measurement systems (adversarial instinct, coverage thinking, regression discipline, documentation).
  • Proficiency in using AI in daily work and building tools to enhance team effectiveness.
  • Sufficient fluency in statistics and experimental design to partner on calibration and reliability.
  • Ability to design workflows and build functional UIs for non-engineers.
  • Genuine interest in improving healthcare and judgment to prioritize issues.
  • Experience in healthcare or another regulated, high-trust domain.
  • Familiarity with the regulatory landscape.
  • Hands-on experience with the eval tooling ecosystem (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar).
  • Red-teaming or AI safety experience (prompt injection, jailbreaks, adversarial and stress testing).
  • Experience with automated, eval-driven model or prompt optimization.
  • Experience building in an early-stage or fast-moving environment.

Skills

  • AI system evaluation
  • LLM-as-judge
  • Red-teaming
  • Adversarial testing
  • Synthetic scenario generation
  • Multi-turn evaluation
  • Agentic evaluation
  • Production systems development
  • Tooling development
  • Statistics
  • Experimental design
  • UI development
  • Prompt injection
  • Jailbreaks
  • Stress testing
  • Model optimization
  • Prompt optimization

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Senior

About the Company

  • Nuna is building an AI health coach to help patients manage chronic conditions at home.
  • The company uses motivational interviewing to help patients and their families design experiments that fit their real lives and navigate the healthcare system.
  • Nuna's mission is to help patients get their lives back.

Equal Opportunity

  • Nuna is an Equal Employment Opportunity employer.
  • All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, disability, genetics and/or veteran status.