Software Engineer, High Performance Computing at Eventual | San Francisco, CA | Rezi

Software Engineer, High Performance Computing at Eventual

Software Engineer, High Performance Computing

Eventual · San Francisco, CA

1 weeks ago

Software Engineer, High Performance Computing

Eventual · San Francisco, CA

11 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a Systems Engineer on the Dataloading team, you will build the layer that turns multi-petabyte video corpora into dict[str, Tensor] directly on the GPU at line rate. You will work with top Physical AI labs training on the newest generation hardware and collaborate with the largest public AI companies on Earth to keep GPUs fed through advanced sampling, caching, and co-loading techniques. This role requires a deep understanding of system performance and byte flow, with opportunities to gain expertise in NVL72, CUDA, and SLURM.

Responsibilities

  • Design and build the video-native dataloader: rank-aware, NVMe-cached, random-access into clips, returns tensors directly to the GPU.
  • Profile and optimize the full data path from object store → NVMe → page cache → host RAM → device RAM, eliminating every avoidable copy and stall.
  • Saturate the latest hardware (B200, GB200, NVL72) on real customer training jobs and push toward Vera Rubin bandwidth requirements.
  • Own performance benchmarks against customer baselines and historical numbers, catching regressions at PR time.
  • Partner with researchers at partner labs to integrate the loader into their training stack and measure MFU end-to-end.
  • Work cross-team with Storage Infrastructure on the index/format boundary and with Visual Understanding on the model-output ingestion path.

Requirements

  • Obsession with systems-level performance; ability to recite Jeff Dean's "numbers every programmer should know" and analyze flamegraphs.
  • Strong opinions on io_uring.
  • Proficiency in Rust, C++, or C.
  • Strong familiarity with operating systems: page cache, scheduling, syscalls, NUMA, memory hierarchies.
  • A sense for where bytes actually go: NVMe vs. memory vs. network vs. PCIe vs. NVLink, and their throughput and latency budgets.
  • Experience working with GPUs is a plus.
  • Experience working with SLURM, Kubernetes for GPU workloads, or other HPC schedulers.
  • Hands-on CUDA experience.
  • Deep expertise on memory and caching subsystems: page cache tuning, hugepages, NUMA pinning, GPU-Direct Storage.
  • Experience with video decode pipelines (PyAV, decord, NVDEC) or PyTorch DataLoader internals.
  • Contributed to open-source systems projects in Rust/C++.

Skills

  • Systems-level performance optimization
  • Rust
  • C++
  • C
  • Operating Systems
  • io_uring
  • NVMe
  • Network
  • Memory Hierarchies
  • PCIe
  • NVLink
  • GPU
  • SLURM
  • Kubernetes
  • HPC Schedulers
  • CUDA
  • Memory and Caching Subsystems
  • Page Cache Tuning
  • Hugepages
  • NUMA Pinning
  • GPU-Direct Storage
  • Video Decode Pipelines
  • PyTorch DataLoader
  • Open-source systems projects

Location

  • San Francisco Mission district office

Work Type

  • In-person
  • 4 days/week

Experience Level

  • Mid-level
  • Senior

Benefits

  • Competitive compensation
  • Meaningful startup equity
  • Catered lunches and dinners
  • Commuter benefit
  • Team-building events
  • Poker nights
  • Health coverage
  • Vision coverage
  • Dental coverage
  • Flexible PTO
  • Latest Apple equipment
  • 401(k) plan with match

About the Company

  • Eventual was founded in 2022 to close the gap in data platforms for AI training.
  • Our open-source engine, Daft, is the distributed data engine purpose-built for multimodal AI.
  • Daft is running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens.
  • We are building a video-native index on top of our engine for Physical AI that streams curated datasets to GPUs at line rate.
  • We are building this in partnership with top Physical AI labs and public AI infrastructure companies.
  • We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and angels from the co-founders of Databricks and Perplexity.
  • We've assembled a world-class team from AWS, Render, Pinecone and Tesla.
  • Our team has spent careers powering the last generation of Physical AI in self-driving and are excited to do this for the next.