About the Role
Our robots are already operating in real homes and businesses, giving us the opportunity to rapidly improve from real-world experience. With a growing team, strong customer demand, and capital for expansion, we’re entering an exciting stage of growth—and we’re looking for people with exceptional talent and standards to help bring home robotics to millions of households. You'll turn raw fleet data into the datasets our models train on, and follow that data into the training loop, shaping decisions on sampling ratios, data mixes, and curricula. The job is equal parts data engineering and data understanding: build the platform that processes millions of episodes and feeds them to training, and know the data well enough to say what correlates with good and bad model behavior.
Responsibilities
- Characterize data coverage, redundancy, and drift using spectral analysis on time-series, distributional statistics, and clustering over embeddings.
- Identify and remove bad data from terabytes of episodes, such as bad trajectories, annotations, dropped frames, and desynced streams, and automate the detection process.
- Build the data lifecycle, including curation, preprocessing, annotation, augmentation, and versioning, along with training-ingest formats and loaders that serve data at full throughput.
- Utilize models as instruments for embedding search to mine scenarios, and use model loss and disagreement as quality signals, with VLM-assisted filtering and labeling.
- Partner with researchers on model failures, build datasets and processing steps to target specific capabilities and failure modes, and own sampling and mixture decisions for training runs.
- Build eval datasets and benchmarks to measure performance across tasks, environments, embodiments, and model versions.
Requirements
- Experience with data engineering at scale, including pipelines over terabytes, object stores (S3, GCS), distributed storage, and indexing.
- Experience with training ingest, including high-throughput formats and dataloaders.
- Working ML experience, including launching fine-tunes, reading loss curves, and designing ablations to test data hypotheses.
- Analytical range, including signal processing and statistics on real sensor data, and unsupervised structure-finding (PCA, UMAP, clustering) when labels are absent.
- Data debugging skills, with the ability to trace problems from sensor drift through corrupted episodes to pipeline failures.
- Strong Python and software engineering fundamentals.
- C++ is a plus.
- Experience with batch processing at scale using Ray, Spark, or Dask over terabytes, with cost and throughput judgment.
- Experience with training ingest formats and dataloaders that keep GPU clusters fed.
- Experience with workflow orchestration in production using Airflow, Kubeflow, or similar, with retries and monitoring.
- Familiarity with robotics data pitfalls, including timestamps, clock domains, and sensor modality quirks.
- Exposure to robot learning, including training policies (VLAs, world models, RL), and the ability to distinguish data problems from model problems.
Skills
- Data Engineering
- MLOps
- Python
- C++
- Signal Processing
- Statistics
- Unsupervised Learning
- Data Debugging
- Ray
- Spark
- Dask
- Airflow
- Kubeflow
- Robot Learning
- VLAs
- World Models
- RL
Location
- Home
- Business
Work Type
- Full-time
Experience Level
- Mid-level
- Senior
About the Company
- Founded to build robots for home use.
- Believes the next generation of robotics will transform everyday life.
- Believes robots must ship to matter.
- Has robots operating in real homes and businesses.
- Is entering an exciting stage of growth with a growing team, strong customer demand, and capital for expansion.
