About the Role
This role blends research with large-scale data engineering to help assemble pre-training datasets and data systems for next-generation AI models. You will design and implement methods for sourcing, curating, and analyzing pre-training data for quality and performance, working with automated pipelines and human-in-the-loop processes. This position is ideal for individuals passionate about the intersection of data, machine learning, and systems, and who are excited by the challenge of shaping frontier AI.
Responsibilities
- Design and implement techniques for curating, sourcing, and filtering large-scale text, code, and multimodal data.
- Develop data quality metrics and analysis to measure coverage, diversity, and representativeness across sources.
- Collaborate with research and infrastructure teams to scale data processing systems efficiently and reproducibly.
- Investigate and mitigate data risks, including privacy, safety, and licensing concerns, to ensure responsible and ethical data use.
- Continuously evaluate dataset improvements by analyzing their downstream effects on model learning and behavior.
- Publish and present research that moves the entire community forward.
- Share code, datasets, and insights that accelerate progress across industry and academia.
Requirements
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
- Comfortable with debugging distributed training and writing code that scales.
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- Clarity in communication, an ability to explain complex technical concepts in writing.
- A strong grasp of probability, statistics, and ML fundamentals.
- Experience with curation, preprocessing, and analysis of large-scale text, code, or multimodal datasets.
- Prior experience in data engineering, dataset construction, or large-scale web data processing for machine learning models.
- Experience evaluating or improving training data quality and knowledge of data ethics, safety, and licensing frameworks relevant to AI dataset creation.
- Contributions to open datasets, research publications, or data tooling.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
Skills
- Python
- Deep learning frameworks (PyTorch, TensorFlow, JAX)
- Debugging distributed training
- Writing scalable code
- Communication
- Probability
- Statistics
- Machine Learning fundamentals
- Data curation
- Data preprocessing
- Data analysis
- Data engineering
- Dataset construction
- Large-scale web data processing
- Data quality evaluation
- Data ethics
- Data safety
- Licensing frameworks for AI dataset creation
- Open datasets
- Research publications
- Data tooling
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Entry-level to Senior
Education Level
- Bachelor's degree or equivalent experience
- PhD or equivalent industry research experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
Equal Opportunity
- We sponsor visas.
