Staff Engineer, Distributed Storage and HPC & AI Infrastructure at Together AI | San Francisco | Rezi

Staff Engineer, Distributed Storage and HPC & AI Infrastructure at Together AI

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AI · San Francisco

1 months ago

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AI · San Francisco

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Design and deliver multi-petabyte storage systems for AI training and inference workloads. Architect high-performance parallel filesystems and object stores, evaluate cutting-edge technologies, and drive cost optimization. Build Kubernetes-native storage operators and self-service platforms for automated provisioning and multi-tenancy. Optimize end-to-end data paths, design multi-tier caching architectures, and implement intelligent prefetching.

Responsibilities

  • Design multi-petabyte AI/ML storage systems; integrate WekaFS, Ceph, etc.; lead capacity planning and cost optimization (30-50% savings via tiering, lifecycle policies, right-sizing).
  • Design/optimize RDMA, InfiniBand, 400GbE networks; tune for max throughput/min latency; implement NVMe-oF/iSCSI; troubleshoot bottlenecks; optimize TCP/IP for storage.
  • Build Kubernetes storage operators/controllers; enable automated provisioning, self-service abstractions, multi-tenant isolation, quotas; create reusable Helm/Terraform patterns.
  • Deliver 10-50 GB/s per GPU node; optimize caching (weights/datasets/checkpoints), parallel filesystems, and data paths; troubleshoot with profiling tools; scale to thousands of nodes.
  • Build multi-tier caches (local NVMe, distributed, object); optimize data locality and model-weight distribution; implement smart prefetching/eviction.
  • Implement monitoring, alerting, SLOs; design DR/backups with runbooks; run chaos engineering; ensure 99.9%+ uptime via proactive/automated remediation.
  • Partner with ML/SRE teams; mentor on storage best practices; contribute to open-source; write docs, postmortems, and public learnings.

Requirements

  • 8+ years in storage engineering with 3+ years managing distributed storage at multi-petabyte scale
  • Proven track record deploying and operating high-performance storage for GPU/HPC clusters
  • Deep Kubernetes and cloud-native storage experience in production environments
  • Strong coding skills in Go and Python with demonstrated ability to build production-grade tools
  • BS/MS in Computer Science, Engineering, or equivalent practical experience
  • History of technical leadership: designing systems that significantly improved performance (>3x), reliability (99.9%+ uptime), or cost efficiency
  • Deep expertise in WekaFS, Lustre, GPFS, BeeGFS, or similar parallel filesystems at multi-petabyte scale
  • Production experience with S3, MinIO, Ceph, or R2 including performance optimization and cost management
  • CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers
  • Storage optimization for GPU workloads, RDMA/InfiniBand networking, parallel filesystem optimization (100+ GB/s aggregate cluster throughput)
  • Go and Python for automation, operators, and tooling
  • Terraform, Ansible, Helm, GitOps (ArgoCD)
  • Advanced knowledge of filesystems (ext4, xfs), LVM, NVMe optimization, RAID configurations
  • Prometheus, Grafana, Thanos architecture and operations

Skills

  • WekaFS
  • Ceph
  • Lustre
  • Go
  • Python
  • Kubernetes
  • RDMA
  • InfiniBand
  • 400GbE
  • NVMe-oF
  • iSCSI
  • TCP/IP
  • Helm
  • Terraform
  • Prometheus
  • Grafana
  • Thanos
  • GPU Direct Storage (GDS)
  • 100GbE
  • Storage snapshots
  • Cloning
  • Thin provisioning
  • Backup and disaster recovery
  • Storage encryption
  • Storage benchmarking
  • Profiling tools

Work Type

  • full-time
  • remote

Experience Level

  • 8+ years in storage engineering
  • 3+ years managing distributed storage at multi-petabyte scale

Education Level

  • BS/MS in Computer Science, Engineering, or equivalent practical experience

Salary/Compensations

  • $250,000 - $300,000

Benefits

  • startup equity
  • health insurance

About the Company

  • Together AI is a research-driven artificial intelligence company.
  • We believe open and transparent AI systems will drive innovation and create the best outcomes for society.
  • We are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
  • We have contributed to leading open-source research, models, and datasets to advance the frontier of AI.
  • Our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama.
  • We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.

Equal Opportunity

  • Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.