Member of Technical Staff - ML Performance at Veeda AI | CA | Rezi

Member of Technical Staff - ML Performance at Veeda AI

Member of Technical Staff - ML Performance

Veeda AI · CA

1 weeks ago

Member of Technical Staff - ML Performance

Veeda AI · CA

10 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities

  • Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core.
  • Converge BF16, FP8, and NVFP4 recipes on Blackwell, addressing scaling-factor and accumulation bugs in video tokenizer and VAE layers.
  • Write and tune CUDA and Triton kernels, driving FlashAttention-4, FlexAttention, and torch.compile integration to optimize attention over long video sequences.
  • Tune NCCL collectives and compute/communication overlap across NVLink domains and fabric, using NCCL flight recorder to diagnose and resolve issues.
  • Build detection layers for silent data corruption, stuck CUDA kernels, and hangs, implementing asynchronous and tiered checkpointing for faster recovery.

Requirements

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.
  • Deep hands-on experience with PyTorch and at least one large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed) on real multi-node jobs.
  • Fluency in Python and C++/CUDA with the ability to predict kernel stalls from memory access patterns.
  • Experience profiling live training runs with Nsight Systems or PyTorch profiler and translating traces into quantifiable improvements.
  • Expertise in low-precision numerics, kernel authoring, or large-run fault diagnosis, with credibility in the others.

Skills

  • PyTorch
  • FSDP2
  • Megatron-Core
  • Python
  • C++
  • CUDA
  • Nsight Systems
  • PyTorch profiler
  • Triton
  • FlashAttention-4
  • FlexAttention
  • torch.compile
  • NCCL
  • NVLink

Education Level

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.

About the Company

  • Veeda AI is building the next generation of multimodal foundation world models for Physical AI.
  • Small, fast-moving team of engineers and researchers from leading AI labs.
  • Tackling challenging problems at the intersection of AI, robotics, and embodied intelligence.
  • Opportunity to make an outsized impact from day one.