About the Role
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. This role involves creating challenging tasks and evaluation criteria within realistic simulated environments to evaluate AI coding agents.
Responsibilities
- Build virtual companies with codebases, infrastructure, and context.
- Assemble and calibrate tasks, including crafting prompts and defining evaluation criteria.
- Design tasks in isolated environments emulating a developer's workstation.
- Write tests that accept all correct solutions and reject incorrect ones.
- Iterate with an AI agent on tests to verify their effectiveness.
- Review code written by agents and design edge cases and adversarial scenarios.
- Iterate based on feedback from expert QA reviewers.
Requirements
- Degree in Computer Science, Software Engineering, or related fields.
- 5+ years in software development, primarily Python (FastAPI, pytest, async/await, subprocess, file operations).
- Background in full-stack development, with experience building React-based interfaces (JavaScript/TypeScript) and robust back-end systems.
- Experience writing functional and integration tests.
- Familiarity with Docker containers and infrastructure tools (Postgres, Kafka, Redis).
- Understanding of CI/CD (GitHub Actions as a user).
- Comfortable reading and reasoning about code across the stack.
Skills
- Python
- FastAPI
- pytest
- async/await
- subprocess
- file operations
- React
- JavaScript
- TypeScript
- Docker
- Postgres
- Kafka
- Redis
- GitHub Actions
Location
- Remote
Work Type
- Project-based
- Part-time
- Non-permanent
Experience Level
- 5+ years in software development
- Experienced developers
- Software engineers
- Test automation specialists
Education Level
- Degree in Computer Science, Software Engineering, or related fields
Salary/Compensations
- Up to $40 per hour equivalent
About the Company
- Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.
