About the Role
Design and build an end-to-end evaluation platform to support AI research, enabling researchers to author, run, and analyze model evaluations efficiently and reliably.
Responsibilities
- Design, build, and maintain the platform for authoring, running, tracking, and analyzing model evaluations.
- Work across evaluation libraries, distributed backend systems, data pipelines, APIs, and user-facing applications.
- Build flexible abstractions for evaluation tasks, environments, graders, datasets, and model outputs.
- Make evaluation results reproducible and trustworthy through versioning, provenance, observability, failure recovery, and robust quality controls.
- Partner directly with researchers to identify bottlenecks and turn bespoke workflows into self-serve systems.
- Work with the Research Tooling engineering team.
Requirements
- A bachelor’s degree, or equivalent practical experience, in computer science, engineering, machine learning, or a related field.
- Two years of post-grad work experience as a software engineer or ML engineer, exclusive of internships.
- Hands-on experience building or maintaining evaluations, benchmarks, graders, or model-quality systems for large language or multimodal models.
- Strong software engineering fundamentals and experience building reliable, maintainable systems.
- Proficiency in at least one backend programming language (Python and Rust primarily used).
- Experience with React and Typescript on the frontend.
- Experience with databases, data pipelines, distributed systems, or other data-intensive infrastructure.
- Comfort working across the stack and owning projects from initial problem discovery through deployment and operation.
- Experience collaborating with cross-functional partners and subject-matter experts.
- A track record of building frameworks, SDKs, or developer tools with thoughtful abstractions and a strong user experience.
- Experience with distributed job execution, workflow orchestration, sandboxed environments, or large-scale data processing.
- Experience building polished, intuitive interfaces for inspecting complex data, comparing experiments, or debugging model behavior.
- Familiarity with large language or multimodal model evaluation.
- Experience working closely with researchers to understand their workflows and turn rapidly evolving needs into durable systems.
- Experience at a startup or on a small team, building technically complex products end to end.
Skills
- Python
- Rust
- React
- Typescript
- Databases
- Data pipelines
- Distributed systems
- Data-intensive infrastructure
- Backend programming
- Frontend development
- Software engineering
- Machine learning
- Large language models
- Multimodal models
- Evaluation frameworks
- Benchmark systems
- Grader systems
- Model quality systems
- Distributed job execution
- Workflow orchestration
- Sandboxed environments
- Large-scale data processing
- User interface design
- Data inspection
- Experiment comparison
- Model behavior debugging
- Model-based grading
- Human evaluation
- Synthetic data generation
Location
- San Francisco, California
Work Type
- Onsite
Experience Level
- Two years of post-grad work experience
Education Level
- Bachelor’s degree or equivalent practical experience
Salary/Compensations
- $300,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
Equal Opportunity
- We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
