Member of Technical Staff - ML Operations at Veeda AI | CA | Rezi

Member of Technical Staff - ML Operations at Veeda AI

Member of Technical Staff - ML Operations

Veeda AI · CA

1 weeks ago

Member of Technical Staff - ML Operations

Veeda AI · CA

10 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities

  • Own how a run is defined, launched, resumed, and killed, from typed configs (Hydra, OmegaConf) to pinned container digests to a relaunch that takes one command.
  • Bind every checkpoint to its code commit, config hash, dataset version, and container digest in Weights & Biases or MLflow, so an old run rebuilds from its manifest, not from memory.
  • Own retention and garbage-collection policy, PyTorch DCP resharding and format conversion, and promotion from raw checkpoint to evaluated artifact that simulation and robotics can safely build on.
  • Gate each checkpoint on seeded rollout and policy-success suites, run per-change and nightly as Slurm arrays under Argo Workflows, with confidence intervals wide enough to separate regression from eval noise.
  • Carry the pager for live runs (loss spikes, throughput cliffs, data loader stalls), and report goodput against allocated GPU-hours as the number capacity decisions actually run on.

Requirements

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.
  • Strong Python and software engineering skills and real CI/CD experience (e.g., GitHub Actions, Buildkite).
  • Shipped internal tooling that other engineers chose to keep using.
  • Operated multi-node training jobs, carried the pager for them, and decided from telemetry whether to kill, requeue, or let a degraded run ride.
  • Built reproducible pipelines end to end, and can precisely state which parts of a training run are bit-reproducible, which are not, and why.
  • Rigorous about evaluation methodology, from seeds and sample sizes to confidence intervals, and can distinguish a genuine regression from a flaky harness.

Skills

  • Python
  • Software Engineering
  • CI/CD
  • GitHub Actions
  • Buildkite
  • Weights & Biases
  • MLflow
  • Hydra
  • OmegaConf
  • Argo Workflows
  • Slurm
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Flyte
  • Ray

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Education Level

  • Bachelor's degree
  • Equivalent hands-on experience

About the Company

  • Veeda AI is building the next generation of multimodal foundation world models for Physical AI.
  • Small, fast-moving team of engineers and researchers from leading AI labs.
  • Tackling challenging problems at the intersection of AI, robotics, and embodied intelligence.