Impress employers and recruiters.
Choose from hundreds of resume examples.

Impress employers and recruiters.
Choose from hundreds of resume examples.
Tailor your resume to this 视频生成模型 · 训练 Infra 工程师 role.
Rezi rewrites your resume against Embedding VC's job description. Free.

Tailor your resume to this 视频生成模型 · 训练 Infra 工程师 role.
Rezi rewrites your resume against Embedding VC's job description. Free.
Don't guess if your resume is good enough.
See how it scores against the 视频生成模型 · 训练 Infra 工程师 posting at Embedding VC — free, in seconds.

Don't guess if your resume is good enough.
See how it scores against the 视频生成模型 · 训练 Infra 工程师 posting at Embedding VC — free, in seconds.
About the Role
We are training our own video generation foundation models (DiT / Flow Matching) and need an engineer who can both build the training platform and transform research code into stable results on hundreds of GPUs. You will be a builder of the platform, not just a user.
Responsibilities
- Build training platforms: From job scheduling, breakpoint retraining, monitoring and alerting, to data/weight pipelines, consolidate scattered scripts into reusable training infrastructure for the team.
- Hundreds of GPUs distributed training: FSDP, Tensor Parallel, Context Parallel, Ulysses, push MFU to a reasonable level.
- PB-level video data pipeline: NVDEC decoding, VAE latent caching, variable resolution bucket sampling.
- GPU memory and performance: FlashAttention, FP8 mixed precision, Triton kernel, activation checkpoint strategy.
- Training stability: Root cause analysis of loss spikes, second-level breakpoint recovery, automatic exclusion of slow nodes.
Requirements
- Proficiency in PyTorch distributed and CUDA architecture.
- Source code level understanding of at least one mainstream training framework (Megatron / DeepSpeed / FSDP / TorchTitan).
- Practical experience with training on ≥ 256 GPUs.
- Experience building training platforms / cluster scheduling / training toolchains from scratch or midway.
Skills
- PyTorch distributed
- CUDA architecture
- Megatron
- DeepSpeed
- FSDP
- TorchTitan
- FSDP
- Tensor Parallel
- Context Parallel
- Ulysses
- NVDEC decoding
- VAE latent caching
- FlashAttention
- FP8 mixed precision
- Triton kernel
- activation checkpoint
- DiT
- Diffusion
- Video data processing
- Triton kernel
- CUTLASS kernel
Experience Level
- ≥ 256 GPUs training practical experience
Benefits
- Access to hundreds of GPUs for real-world training.
- A team that treats training as an engineering problem.
- Opportunities for open-sourcing and publishing within compliance boundaries.