Research, Pre-Training Data at Thinking Machines Lab | CA, US | Rezi

Research, Pre-Training Data at Thinking Machines Lab

Research, Pre-Training Data

Thinking Machines Lab · CA, US

3 weeks ago

Research, Pre-Training Data

Thinking Machines Lab · CA, US

24 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

This role blends research with large-scale data engineering to help assemble pre-training datasets and data systems for next-generation AI models. You will design and implement methods for sourcing, curating, and analyzing pre-training data for quality and performance, working with automated pipelines and human-in-the-loop processes. This position is ideal for individuals passionate about the intersection of data, machine learning, and systems, and who are excited by the challenge of shaping frontier AI.

Responsibilities

  • Design and implement techniques for curating, sourcing, and filtering large-scale text, code, and multimodal data.
  • Develop data quality metrics and analysis to measure coverage, diversity, and representativeness across sources.
  • Collaborate with research and infrastructure teams to scale data processing systems efficiently and reproducibly.
  • Investigate and mitigate data risks, including privacy, safety, and licensing concerns, to ensure responsible and ethical data use.
  • Continuously evaluate dataset improvements by analyzing their downstream effects on model learning and behavior.
  • Publish and present research that moves the entire community forward.
  • Share code, datasets, and insights that accelerate progress across industry and academia.

Requirements

  • Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
  • Comfortable with debugging distributed training and writing code that scales.
  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
  • Clarity in communication, an ability to explain complex technical concepts in writing.
  • A strong grasp of probability, statistics, and ML fundamentals.
  • Experience with curation, preprocessing, and analysis of large-scale text, code, or multimodal datasets.
  • Prior experience in data engineering, dataset construction, or large-scale web data processing for machine learning models.
  • Experience evaluating or improving training data quality and knowledge of data ethics, safety, and licensing frameworks relevant to AI dataset creation.
  • Contributions to open datasets, research publications, or data tooling.
  • PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.

Skills

  • Python
  • Deep learning frameworks (PyTorch, TensorFlow, JAX)
  • Debugging distributed training
  • Writing scalable code
  • Communication
  • Probability
  • Statistics
  • Machine Learning fundamentals
  • Data curation
  • Data preprocessing
  • Data analysis
  • Data engineering
  • Dataset construction
  • Large-scale web data processing
  • Data quality evaluation
  • Data ethics
  • Data safety
  • Licensing frameworks for AI dataset creation
  • Open datasets
  • Research publications
  • Data tooling

Location

  • San Francisco, California

Work Type

  • Full-time

Experience Level

  • Entry-level to Senior

Education Level

  • Bachelor's degree or equivalent experience
  • PhD or equivalent industry research experience

Salary/Compensations

  • $350,000 - $475,000 USD

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support

About the Company

  • The mission of Thinking Machines is to build AI that extends human will and judgment.

Equal Opportunity

  • We sponsor visas.