About the Role
XDOF is building the foundation for foundation models in robotics, focusing on data collection systems, annotation pipelines, exabyte-scale data infrastructure, and software toolchains. This role involves leading technical efforts at the intersection of vision-language models and robot learning, creating systems to transform video data into high-signal training data for VLA models and contributing to model development. The role also drives research into data utility, exploring new metadata and structured annotations to enhance robot learning capabilities.
Responsibilities
- Design and implement vision-language pipelines for egocentric and teleoperation video, including structured captioning, temporal grounding, action-conditioned scene understanding, and semantic annotation at scale.
- Develop and evaluate representations that bridge visual perception, language, and low-level robot action, spanning VLAs, video prediction, and world models.
- Build and improve data curation systems to assess quality, diversity, and coverage of large-scale robot demonstration datasets.
- Work hands-on with bimanual and high-DoF manipulation data, including real teleoperation footage and sim-generated rollouts.
- Collaborate with partner labs to define data requirements and link data quality to downstream policy performance.
- Stay current on research frontiers (VLAs, video foundation models, flow matching, DiT architectures, egocentric pretraining) and translate insights into production systems.
Requirements
- MS or PhD in Computer Science, Robotics, Machine Learning, or a related field from a top-tier program.
- 3–7 years of research or applied research experience (industry or academic) in one or more of: vision-language models, video understanding, robot learning, or generative modeling.
- Deep fluency in PyTorch.
- Working knowledge of large-scale training infrastructure (distributed training, mixed precision, large batch workflows).
- Published work or demonstrable impact in VLMs/VLAs, video representation learning, imitation learning, or a closely related area.
- Strong engineering fundamentals; ability to design clean systems, not just run experiments.
Skills
- Vision-language models
- Video understanding
- Robot learning
- Generative modeling
- PyTorch
- Large-scale training infrastructure
- VLMs/VLAs
- Video representation learning
- Imitation learning
Location
- San Mateo
Work Type
- Flexible work arrangements
Experience Level
- Mid Level to Senior Research Scientist (L4–L5 equivalent)
- Junior candidates will still be considered
Education Level
- MS or PhD in Computer Science, Robotics, Machine Learning, or a related field
Benefits
- Competitive compensation and equity
- Comprehensive health and wellness benefits
- Flexible work arrangements
- Collaborative and fast-paced work environment
- Opportunity to shape the future of robotics and AI alongside an ambitious, values-driven team
About the Company
- Frontier labs are racing to build general-purpose robots, and the bottleneck isn't compute. It's data.
- At XDOF, we're building the foundation behind the foundation models: the data collection systems, annotation pipelines, exabyte-scale data infrastructure, and software toolchain that enable our partners to push the field forward.
