About the Role
At XDOF, we are at an inflection point where high-quality training data is the bottleneck for building general-purpose robots. We are developing the foundational data collection systems, operational capabilities, exabyte-scale data warehouse, and software toolchain to support our partners in advancing the field. The Perception Algorithm team transforms raw multimodal sensor data into high-quality robot training annotations. This role involves the complete loop from data collection to model delivery, including sensor calibration, SLAM localization, human pose estimation, perception model training, and embedded deployment, directly impacting training data quality.
Responsibilities
- Design and optimize hand pose estimation pipelines for accurate joint angle extraction from teleoperation data.
- Build full-body pose estimation systems for motion capture and teleoperation action annotation ground truth generation.
- Research and apply markerless vision-based pose estimation methods to reduce data collection costs.
- Fuse pose estimation outputs with robot joint angle data to generate consistent training annotations.
- Design and maintain intrinsic/extrinsic calibration pipelines for multi-camera arrays (factory and online recalibration).
- Build visual SLAM / V-SLAM systems for real-time localization and scene reconstruction on data collection platforms.
- Implement hand-eye calibration between cameras and robot end-effectors.
- Develop temporal alignment solutions across multimodal sensors (cameras, IMU, data gloves, force sensors).
- Train and iterate on perception models, including object detection, instance segmentation, and 6DoF pose estimation.
- Optimize model inference using TensorRT / CUDA for real-time performance on robot embedded platforms.
- Write custom CUDA kernels for low-level acceleration of perception tasks.
- Design evaluation metric frameworks for perception models and continuously track the relationship between model performance and data quality.
- Contribute to the design of automated annotation pipelines that convert sensor data into structured training labels.
- Build Auto QA modules to filter low-quality data, including anomalous frames, failed demonstrations, and sensor dropouts.
- Collaborate with ML engineers and data infrastructure teams to ensure perception output formats meet downstream VLA model training requirements.
- Establish feedback mechanisms linking perception accuracy to model training outcomes for continuous annotation quality improvement.
Requirements
- 5+ years of industry experience in robot perception or computer vision.
- Strong 3D vision fundamentals: stereo and structured-light camera principles, 3D reconstruction.
- Proficiency with SLAM frameworks (ORB-SLAM, VINS-Mono, FastLIO, etc.) or V-SLAM system development experience.
- Hands-on engineering experience with human pose estimation: hand joints (MediaPipe, MANO) or full-body pose (OpenPose, SMPLify, etc.).
- Proficient in deep learning training frameworks for perception model training, tuning, and evaluation.
- TensorRT deployment experience with real-time inference optimization on embedded platforms (Jetson, Horizon, etc.).
- CUDA programming fundamentals; ability to write or debug custom kernels.
- Proficient in C++ and Python with ROS / ROS2 development experience.
- Proficient with AI coding agents.
- Engineering experience with 6DoF object pose estimation (FoundPose, FoundationPose, GDR-Net, etc.).
- Familiarity with 3D Gaussian Splatting or NeRF for scene reconstruction or data augmentation.
- Experience with robot manipulation or teleoperation systems.
- End-to-end development experience with automated annotation pipelines or ground truth generation systems.
- Published research in perception, pose estimation, or robotics.
Skills
- Robot perception
- Computer vision
- 3D vision fundamentals
- Stereo camera principles
- Structured-light camera principles
- 3D reconstruction
- SLAM frameworks
- V-SLAM system development
- Human pose estimation
- Deep learning training frameworks
- TensorRT
- CUDA programming
- C++
- Python
- ROS
- ROS2
- AI coding agents
- 6DoF object pose estimation
- 3D Gaussian Splatting
- NeRF
- Robot manipulation
- Teleoperation systems
- Automated annotation pipelines
- Ground truth generation systems
- Research in perception, pose estimation, or robotics
Experience Level
- 5+ years of industry experience
Benefits
- Direct involvement in the most critical technical challenge in embodied intelligence: producing high-quality robot training data.
- An environment working alongside top-tier robotics engineers and ML researchers.
- Access to proprietary hardware platforms (humanoid robots, camera arrays, data gloves).
- A fast-paced, high-autonomy 0→1 work environment.
About the Company
- XDOF is at an inflection point, addressing the bottleneck of high-quality training data for general-purpose robots.
- The company is building the foundation behind foundation models, including data collection systems, operational capability, an exabyte-scale data warehouse, and a software toolchain.
- XDOF aims to help partners drive the field of embodied intelligence forward.
