Member of the Technical Staff - Platform at Andromeda Cluster | San Francisco, California, United States | Rezi

Member of the Technical Staff - Platform at Andromeda Cluster

Member of the Technical Staff - Platform

Andromeda Cluster · San Francisco, California, United States

4 days ago

Member of the Technical Staff - Platform

Andromeda Cluster · San Francisco, California, United States

5 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are looking for engineers to build and operate the control plane that runs our fleet.

Responsibilities

  • Build and operate the control plane for the fleet.
  • Develop automated systems to provision clusters from bare machines to customer-ready.
  • Manage machine lifecycle between tenants, including joining, wiping, verifying, and rejoining.
  • Operate Kubernetes and Postgres across the fleet.
  • Contribute to custom Kubernetes operators.
  • Scale clusters from tens to thousands of nodes.
  • Participate in on-call rotations.

Requirements

  • Demonstrate impressive technical work with real-world impact.
  • Possess 2+ years of on-call experience for critical production services.
  • Have deep Kubernetes experience.
  • Exhibit strong Linux fundamentals, including kernel, cgroups, containers, networking, and storage.
  • Have experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale.
  • Be familiar with fleet management and capacity planning.
  • Prior experience writing operators, or with GPUs, HPC scheduling, or bare-metal hardware is a plus.

Skills

  • Kubernetes
  • Postgres
  • Linux
  • Databases
  • Monitoring
  • CI/CD
  • Cloud Infrastructure
  • Fleet Management
  • Capacity Planning
  • Operators
  • GPUs
  • HPC Scheduling
  • Bare-metal Hardware

Location

  • North America Remote
  • San Francisco, CA

Work Type

  • Full-Time

Experience Level

  • 2+ years of on-call experience

About the Company

  • Andromeda is building the liquidity layer for compute, enabling startups to access scaled AI infrastructure.
  • Founded by Nat Friedman and Daniel Gross, Andromeda provides a platform that deploys into foreign datacenters and turns hardware into clusters for AI training.
  • The company operates compute for over 80 customers across 30+ capacity providers, managing tens of thousands of GPUs and supporting billions of GPU-hours.

Equal Opportunity

  • Andromeda Cluster is an equal opportunity employer.
  • We celebrate diversity and are committed to creating an inclusive environment for all employees.
  • We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.