Member of Technical Staff - Data at Veeda AI | CA | Rezi

Member of Technical Staff - Data at Veeda AI

Member of Technical Staff - Data

Veeda AI · CA

1 weeks ago

Member of Technical Staff - Data

Veeda AI · CA

10 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities

  • Build ingest for video, lidar, and robot trajectories on Ray Data and Daft, with GPU decode (NVDEC, DALI) and resharding into WebDataset and Lance layouts that stream sequentially rather than seeking per sample.
  • Decide what earns a slot using blur, exposure, and camera-trajectory scoring plus embedding deduplication over cuVS indexes, and prove each filter with a downstream ablation, not a dataset-size delta.
  • Produce the labels the models need, such as VLM captions, camera pose from feed-forward reconstruction (VGGT, MASt3R), and depth and segmentation pseudo-labels, and hold each to a measured error rate against human review.
  • Normalize episodic data across formats such as LeRobotDataset v3, Open X-Embodiment, and RLDS, reconciling action spaces, control rates, and frame timing, and account for the simulated share of every training mixture.
  • Track license terms, restricted-source flags, and C2PA content credentials at source granularity, and version datasets as immutable manifests so any checkpoint traces back to the exact bytes that trained it.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in large-scale data engineering.
  • Built and operated distributed data pipelines (e.g., Ray Data, Daft, Spark) over hundreds of terabytes, with rigor in idempotency, backfills, and schema evolution.
  • Strong Python skills and comfortable in the video stack (codecs, containers, ffmpeg, GPU decode) and in columnar and object storage formats.
  • Able to design and defend a data mixture empirically, running curation ablations that measure downstream model quality.
  • Experience working inside real licensing constraints on what may and may not be trained on, with provenance treated as a hard requirement.

Skills

  • Python
  • Video stack (codecs, containers, ffmpeg, GPU decode)
  • Columnar and object storage formats
  • Ray Data
  • Daft
  • Spark
  • Robot trajectory formats
  • LeRobot
  • RLDS
  • ROS 2 bags
  • MCAP
  • Sensor calibration
  • Hardware time synchronization
  • Non-pinhole camera models (fisheye, ftheta)
  • GPU-accelerated curation
  • RAPIDS
  • NeMo Curator
  • PII, face, and plate redaction
  • Annotation vendor management
  • QA statistics
  • Lakehouse storage
  • Iceberg
  • Delta
  • Data curation
  • Open-source data tooling
  • DataTrove
  • video2dataset

Education Level

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in large-scale data engineering.

About the Company

  • Veeda AI is building the next generation of multimodal foundation world models for Physical AI.
  • We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence.
  • If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.