Member of Technical Staff — Compute Cluster at Causal Labs | San Francisco, USA | Rezi

Member of Technical Staff — Compute Cluster at Causal Labs

Member of Technical Staff — Compute Cluster

Causal Labs · San Francisco, USA

1 weeks ago

Member of Technical Staff — Compute Cluster

Causal Labs · San Francisco, USA

12 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are building a Large Physics foundation Model (LPM) to achieve general causal intelligence, capable of predicting the future and identifying actions to alter it. This role is critical in designing, building, and operating the supercomputing environment that powers our GPU fleet, enabling rapid research iteration at scale.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters end to end, including provisioning, imaging, upgrades, and capacity planning.
  • Extend scheduling and orchestration systems (e.g., Kubernetes, Slurm) for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads.
  • Build software to abstract cluster management and provide a unified, self-serve interface to researchers and engineers.
  • Own cluster storage and artifact paths for checkpoints and logs, ensuring clear retention and lineage.
  • Monitor and continuously improve reliability and error recovery, building observability to proactively identify failures.
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs.

Requirements

  • Experience operating large-scale GPU clusters and container orchestration frameworks (e.g., Kubernetes, Slurm, Docker).
  • Strong systems background including Linux, networking, storage, and infrastructure-as-code.
  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings.
  • Understanding of monitoring, logging, observability, and version control best practices for ML systems.
  • Familiarity with CUDA/NCCL and performance profiling for distributed workloads.
  • Ability to own deliverables end-to-end, from requirements through autonomous execution.

Skills

  • GPU cluster operation
  • Container orchestration (Kubernetes, Slurm, Docker)
  • Linux
  • Networking
  • Storage
  • Infrastructure-as-code
  • Cloud platforms (GCP, AWS, Azure)
  • ML/AI services
  • Monitoring
  • Logging
  • Observability
  • Version control
  • CUDA
  • NCCL
  • Performance profiling

About the Company

  • Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it.
  • We are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect.
  • We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.
  • Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.