Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware) at Veeda AI | CA | Rezi

Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware) at Veeda AI

Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Veeda AI · CA

1 weeks ago

Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Veeda AI · CA

10 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities

  • Design, optimize, and maintain high-throughput distributed training systems across large-scale GPU clusters for multi-modal foundation models.
  • Debug, diagnose, and resolve subtle numerical instability issues in FP16, BF16, FP8, and custom quantization schemes.
  • Build advanced fault-detection mechanisms and automated diagnostics to rapidly pinpoint and isolate silent data corruption, hardware hang/deadlock, memory leaks, and "card-freeze" issues during large training runs.
  • Profile distributed communication bottlenecks, memory usage, and kernel execution to improve overall FLOPS utilization across multi-node, multi-GPU training jobs.
  • Develop resilient checkpointing systems, rapid fault-recovery pipelines, and execution telemetry to keep researcher productivity high and hardware downtime minimal.

Requirements

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.
  • Deep hands-on experience with deep learning training frameworks (e.g., PyTorch) and distributed training paradigms (FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism).
  • Proven experience in numerical precision analysis, low-precision training (BF16/FP8), and debugging complex loss divergence/stability issues in massive training runs.
  • Strong root-cause analysis skills for hardware/software interaction bugs, including stuck CUDA kernels, NCCL timeouts, GPU hardware faults, and silent training corruptions.
  • Strong programming skills in Python and C++/CUDA, with a deep understanding of low-level GPU architectures and memory hierarchies.

Skills

  • PyTorch
  • FSDP
  • Megatron-LM
  • DeepSpeed
  • Tensor Parallelism
  • Pipeline Parallelism
  • BF16
  • FP8
  • Python
  • C++
  • CUDA

Education Level

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.

About the Company

  • Veeda AI is building the next generation of multimodal foundation world models for Physical AI.
  • We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence.