About the Role
This role advances the science of how large models learn from data by exploring new pre-training methods, architectures, and learning objectives. It blends fundamental research and practical engineering, requiring high-performance coding and deep theoretical exploration.
Responsibilities
- Research and develop new methodologies for pre-training.
- Work in areas such as scaling, architecture, algorithms, or optimization of large scale training runs.
- Design data curricula and sampling strategies that improve learning dynamics and model generalization.
- Collaborate with infrastructure and data teams to conduct large-scale experiments efficiently and reproducibly.
- Publish and present research that moves the entire community forward.
- Share code, datasets, and insights that accelerate progress across industry and academia.
Requirements
- Ability to design, run, and analyze experiments thoughtfully, with demonstrated research judgment and empirical rigor.
- Experience with distributed or high-performance computing environments.
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
- Comfortable with debugging distributed training and writing code that scales.
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- Clarity in communication, an ability to explain complex technical concepts in writing.
- A strong grasp of probability, statistics, and ML fundamentals.
- Prior experience training or analyzing large-scale models, or contributing to pre-training or foundation model research.
- Strong publication record or open-source contributions in representation learning, optimization, scaling laws, or other areas of pre-training.
- Familiarity with curriculum learning, data selection, or active learning techniques.
- Experience designing or maintaining evaluation frameworks for large models.
- Contributions to open datasets, research publications, or data tooling.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
Skills
- Python
- PyTorch
- TensorFlow
- JAX
- Deep Learning
- Distributed Training
- High-Performance Computing
- Probability
- Statistics
- Machine Learning Fundamentals
- Representation Learning
- Optimization
- Scaling Laws
- Curriculum Learning
- Data Selection
- Active Learning
- Evaluation Frameworks
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Entry-level to Senior
Education Level
- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline; or, equivalent industry research experience
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
Equal Opportunity
- We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
