About the Role
Thinking Machines builds multimodal-first AI that extends human will and judgment. We are seeking individuals to advance the science of visual perception and multimodal learning, focusing on the interaction between vision and language at scale. This role involves designing architectures that fuse pixels and text, building datasets and evaluation methods for real-world comprehension, and developing representations that ground abstract concepts in the physical world, aiming for seamless integration into real-world environments. The position blends fundamental research and practical engineering, requiring high-performance code and the ability to read technical reports.
Responsibilities
- Own research projects on training and performance analysis of multimodal AI models.
- Curate and build large-scale datasets and evaluation benchmarks to advance vision capabilities.
- Work with data infrastructure engineers, pretraining researchers and engineers, and the product team to create frontier multimodal models and leverage them in products.
- Publish and present research that moves the entire community forward.
- Share code, datasets, and insights that accelerate progress across industry and academia.
Requirements
- Ability to design, run, and analyze experiments thoughtfully, with demonstrated research judgment and empirical rigor.
- Understanding of machine learning fundamentals, large-scale training, and distributed compute environments.
- Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX).
- Comfortable with debugging distributed training and writing code that scales.
- Clarity in communication, an ability to explain complex technical concepts in writing.
- Research or engineering contributions in visual reasoning, spatial understanding, or multimodal architecture design.
- Experience developing evaluation frameworks for multimodal tasks.
- Publications or open-source contributions in vision-language modeling, video understanding, or multimodal AI.
- A strong grasp of probability, statistics, and ML fundamentals.
- Ability to look at experimental data and distinguish between real effects, noise, and bugs.
Skills
- Python
- PyTorch
- TensorFlow
- JAX
- Machine Learning
- Large-scale training
- Distributed compute
- Visual reasoning
- Spatial understanding
- Multimodal architecture design
- Probability
- Statistics
Location
- San Francisco, California
Work Type
- Full-time
Experience Level
- Bachelor's degree or equivalent experience
- PhD or equivalent industry research experience
Education Level
- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.
Salary/Compensations
- $350,000 - $475,000 USD
Benefits
- Generous health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
About the Company
- The mission of Thinking Machines is to build AI that extends human will and judgment.
- Thinking Machines builds multimodal-first.
Equal Opportunity
- We sponsor visas.
- While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
