Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this Software Engineer, AI Evaluation role.
Rezi rewrites your resume against Nuna's job description. Free.

Tailor your resume to this Software Engineer, AI Evaluation role.
Rezi rewrites your resume against Nuna's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the Software Engineer, AI Evaluation posting at Nuna — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the Software Engineer, AI Evaluation posting at Nuna — free, in seconds.
About the Role
Nuna is building an AI health coach to help patients manage chronic conditions at home. This role involves owning the end-to-end evaluation system for AI agents, ensuring they are safe and effective before deployment.
Responsibilities
- Build testing harnesses and evaluation infrastructure for agentic products and internal tooling.
- Own the architecture and content of AI agent evaluations end-to-end.
- Ensure all agentic deployments pass through the testing apparatus before shipping.
- Own release gates to prevent unsafe or low-quality behavior from reaching patients.
- Build ground truth, judges, and metrics, validating evaluation trustworthiness.
- Build functional tooling for labeling and review workflows for non-engineers.
- Facilitate the loop from evaluation results to model and prompt refinement for safe iteration.
Requirements
- Significant experience building and shipping reliable production systems and tooling.
- Deep understanding of AI system evaluation methods (LLM-as-judge, red-teaming, synthetic scenario generation, multi-turn and agentic evaluation).
- Experience deploying production evals and automated AI tooling.
- A testing mindset applied to building measurement systems (adversarial instinct, coverage thinking, regression discipline, documentation).
- Proficiency in using AI in daily work and building tools to enhance team effectiveness.
- Sufficient fluency in statistics and experimental design to partner on calibration and reliability.
- Ability to design workflows and build functional UIs for non-engineers.
- Genuine interest in improving healthcare and judgment to prioritize issues.
- Experience in healthcare or another regulated, high-trust domain.
- Familiarity with the regulatory landscape.
- Hands-on experience with the eval tooling ecosystem (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar).
- Red-teaming or AI safety experience (prompt injection, jailbreaks, adversarial and stress testing).
- Experience with automated, eval-driven model or prompt optimization.
- Experience building in an early-stage or fast-moving environment.
Skills
- AI system evaluation
- LLM-as-judge
- Red-teaming
- Adversarial testing
- Synthetic scenario generation
- Multi-turn evaluation
- Agentic evaluation
- Production systems development
- Tooling development
- Statistics
- Experimental design
- UI development
- Prompt injection
- Jailbreaks
- Stress testing
- Model optimization
- Prompt optimization
Location
- Remote
Work Type
- Full-time
Experience Level
- Senior
About the Company
- Nuna is building an AI health coach to help patients manage chronic conditions at home.
- The company uses motivational interviewing to help patients and their families design experiments that fit their real lives and navigate the healthcare system.
- Nuna's mission is to help patients get their lives back.
Equal Opportunity
- Nuna is an Equal Employment Opportunity employer.
- All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, disability, genetics and/or veteran status.