视频生成模型 · 训练 Infra 工程师 at Embedding VC | CA, US | Rezi

视频生成模型 · 训练 Infra 工程师 at Embedding VC

视频生成模型 · 训练 Infra 工程师

Embedding VC · CA, US

1 months ago

视频生成模型 · 训练 Infra 工程师

Embedding VC · CA, US

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this 视频生成模型 · 训练 Infra 工程师 role.

Rezi rewrites your resume against Embedding VC's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the 视频生成模型 · 训练 Infra 工程师 posting at Embedding VC — free, in seconds.

About the Role

We are training our own video generation foundation models (DiT / Flow Matching) and need an engineer who can both build the training platform and transform research code into stable results on hundreds of GPUs. You will be a builder of the platform, not just a user.

Responsibilities

  • Build training platforms: From job scheduling, breakpoint retraining, monitoring and alerting, to data/weight pipelines, consolidate scattered scripts into reusable training infrastructure for the team.
  • Hundreds of GPUs distributed training: FSDP, Tensor Parallel, Context Parallel, Ulysses, push MFU to a reasonable level.
  • PB-level video data pipeline: NVDEC decoding, VAE latent caching, variable resolution bucket sampling.
  • GPU memory and performance: FlashAttention, FP8 mixed precision, Triton kernel, activation checkpoint strategy.
  • Training stability: Root cause analysis of loss spikes, second-level breakpoint recovery, automatic exclusion of slow nodes.

Requirements

  • Proficiency in PyTorch distributed and CUDA architecture.
  • Source code level understanding of at least one mainstream training framework (Megatron / DeepSpeed / FSDP / TorchTitan).
  • Practical experience with training on ≥ 256 GPUs.
  • Experience building training platforms / cluster scheduling / training toolchains from scratch or midway.

Skills

  • PyTorch distributed
  • CUDA architecture
  • Megatron
  • DeepSpeed
  • FSDP
  • TorchTitan
  • FSDP
  • Tensor Parallel
  • Context Parallel
  • Ulysses
  • NVDEC decoding
  • VAE latent caching
  • FlashAttention
  • FP8 mixed precision
  • Triton kernel
  • activation checkpoint
  • DiT
  • Diffusion
  • Video data processing
  • Triton kernel
  • CUTLASS kernel

Experience Level

  • ≥ 256 GPUs training practical experience

Benefits

  • Access to hundreds of GPUs for real-world training.
  • A team that treats training as an engineering problem.
  • Opportunities for open-sourcing and publishing within compliance boundaries.