About the Role
Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.
Responsibilities
- Architect and own the scraping and crawling infrastructure that acquires petabytes of image and video data.
- Design systems that are resilient, compliant, and able to keep pace with a fast-moving content landscape.
- Build and maintain production data pipelines that clean, transform, and process raw visual assets.
- Handle format diversity, resolution variance, metadata extraction, and noise at scale.
- Design ETL/ELT workflows that prepare image and video datasets for ML model training and evaluation.
- Include annotation ingestion, quality scoring, and reproducible dataset versioning.
- Continuously profile and optimize pipelines for throughput, latency, and cloud storage/compute costs.
- Ensure research velocity without unnecessary overhead.
- Select, deploy, and operate distributed data-processing frameworks (e.g., Ray, Spark, Dask) to handle compute-intensive tasks.
- Work directly with Computer Vision Engineers, ML Scientists, and product teams to translate data requirements into reliable, well-documented pipelines.
- Close the feedback loop between model performance and data quality.
- Lead data-pipeline design reviews.
- Mentor junior and mid-level data engineers.
- Establish best practices for the data team.
- Raise the overall engineering bar.
Requirements
- 5+ years of professional experience in data engineering or a closely related discipline.
- Strong foundation in distributed systems, database design, and handling large binary assets (images, video) at scale.
- Strong programming skills in Python and SQL.
- Proven experience designing end-to-end production systems rather than basic scripts.
- Designed and operated web scraping or data ingestion systems running reliably in production at significant global scale.
- Hands-on with data-processing frameworks (e.g., PySpark, Pandas, Ray).
- Comfortable with image-processing libraries (OpenCV, Pillow, scikit-image).
- Built ETL/ELT pipelines for ML/CV use cases with a strong emphasis on dataset provenance, lineage, and reproducibility.
- Extensive experience with cloud object storage (AWS S3, GCP GCS).
- Understand the cost, latency, and durability trade-offs when managing petabyte-scale data assets.
- Daily workflow runs through AI coding harnesses (e.g., AI agents/assistants) without sacrificing software engineering rigor.
- Review agent code diffs with the same scrutiny as a team member's PR.
- Recognize AI code generation failure modes.
- Ship rapidly without introducing technical debt or "slop."
Skills
- Python
- SQL
- Distributed systems
- Database design
- Large binary asset handling
- Web scraping
- Data ingestion
- PySpark
- Pandas
- Ray
- OpenCV
- Pillow
- scikit-image
- ETL/ELT pipelines
- ML/CV use cases
- Dataset provenance
- Dataset lineage
- Dataset reproducibility
- AWS S3
- GCP GCS
- AI coding harnesses
- Airflow
- Dagster
- Prefect
- Computer Vision
- Image annotation formats
- MLOps
- Data versioning
- DVC
- LakeFS
- Experiment tracking
- Docker
- Kubernetes
- Data compliance
- GDPR
- Content licensing
- Image rights
Location
- Remote
Work Type
- Full-time
Experience Level
- 5+ years of professional experience
About the Company
- Veeda AI is building the next generation of multimodal foundation world models for Physical AI.
- We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence.
