ML Research Engineer, Training at Weave Robotics | CA, US | Rezi

ML Research Engineer, Training at Weave Robotics

ML Research Engineer, Training

Weave Robotics · CA, US

Yesterday

ML Research Engineer, Training

Weave Robotics · CA, US

a day ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Most robot learning research is graded on evals that don't survive contact with the field. Ours is graded by robots doing useful work in real homes and businesses, every day. We're one of the first companies with a deployed fleet generating real-world robot data at terabyte scale. The pipeline and training stack you build are what turns that data into capability. Model quality is set as much by training as by architecture: what data gets in, how it's sampled, whether the run is stable, whether a silent bug ate the gradient three days ago. You'll own that layer from raw fleet uploads to the batch that hits the GPU. When the stack is right, ideas become models in training in days, and deployed in weeks.

Responsibilities

  • Build training stack end to end: distributed training, data loading, checkpointing, run orchestration, experiment tracking.
  • Develop high-throughput data ingestion, transformation, and storage systems capable of processing terabytes of multimodal robot data, including video, proprioception, and sensor streams.
  • Grow the codebase that makes experiments reproducible, scalable, and easy to launch, monitor, retry/recover, and debug.
  • Drive sampling and curation decisions that show up in model behavior.
  • Profile and fix throughput bottlenecks, chase down loss spikes and silent data bugs, keep results reproducible enough to trust: from data loading to GPU kernels.
  • Turn research prototypes into infrastructure the whole team trains on.

Requirements

  • Deep PyTorch or JAX experience, including multi-node distributed training (FSDP, DDP, or equivalent) on real workloads.
  • Experience profiling and optimizing GPU utilization, data pipelines, I/O bottlenecks, memory usage, and distributed training performance, including CUDA-level profiling tools (e.g. Nsight Systems) and NCCL tuning.
  • You can read a loss curve, tell instability from a data bug and know when to kill a run.

Skills

  • PyTorch
  • JAX
  • Multi-node distributed training
  • FSDP
  • DDP
  • Performance engineering
  • GPU utilization optimization
  • Data pipeline optimization
  • I/O bottleneck optimization
  • Memory usage optimization
  • Distributed training performance optimization
  • CUDA-level profiling tools
  • Nsight Systems
  • NCCL tuning
  • Robot learning
  • Training policies
  • VLAs
  • World models
  • RL
  • Cluster and cloud infrastructure
  • Kubernetes
  • SLURM
  • GCP
  • AWS
  • Large-scale post-training
  • SFT
  • Reward modeling
  • RL fine-tuning
  • On-robot inference optimization
  • TensorRT
  • Quantization
  • Distillation
  • CUDA kernel work
  • Triton kernel work
  • Video-heavy datasets
  • Transcoding
  • Chunking
  • Storage/compute tradeoffs for video training

About the Company

  • Weave was founded to build the robots we’d want to have in our own home.
  • We believe the next generation of robotics will transform everyday life by enabling people to do more and to reclaim time to spend on what’s important.
  • We also believe robots are in a sense like any other product: to matter, they have to ship.
  • Our robots are already operating in real homes and businesses, giving us the opportunity to rapidly improve from real-world experience.
  • With a growing team, strong customer demand, and capital for expansion, we’re entering an exciting stage of growth—and we’re looking for people with exceptional talent and standards to help bring home robotics to millions of households.