Staff Software Engineer, GPU Infrastructure Lifecycle Management at Together AI | United States | Rezi

Staff Software Engineer, GPU Infrastructure Lifecycle Management at Together AI

Staff Software Engineer, GPU Infrastructure Lifecycle Management

Together AI · United States

2 weeks ago

Staff Software Engineer, GPU Infrastructure Lifecycle Management

Together AI · United States

16 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Software Engineer to develop systems that treat infrastructure as software. This role involves owning the software state machines responsible for provisioning hardware, bringing it into service, and managing its entire lifecycle. The goal is to enable the Research and Inference team to provision, scale, and tear down clusters via a single API call, eliminating manual intervention. The platform is manifest-driven, requiring the system to reconcile reality with declared desired states continuously throughout the hardware and cluster lifecycle. You will design the engines that manifest the schema, execute against it, and manage the workflows for hardware and cluster state transitions, from bare metal to a fully functioning AI cluster.

Responsibilities

  • Build the provisioning state machine: design and implement software modeling the full lifecycle of a physical host, including discovery, inference bring-up, GPU driver/CUDA stack, health validation, and decommission/RMA, as explicit, versioned states and transitions.
  • Build the self-service API: design declarative APIs and a control plane for the inference team to request, scale, and tear down inference clusters with a single API call.
  • Automate self-healing: detect and safely drain degraded or failed nodes, trigger repair or replacement, and automatically reintroduce healthy capacity.
  • Own pipeline reliability: ensure idempotency, retries, rollback, and drift detection for a dependable provisioning system.
  • Partner with the inference/ML platform team: understand and encode necessary cluster shapes (topology, interconnect, scheduling constraints) as first-class abstractions.
  • Engineer it like software: apply strong typing, automated tests, code review, versioning, and CI/CD to infrastructure code, treating it as a product.

Requirements

  • Strong software engineering background in Go, Python, Rust, or similar.
  • Experience with durable workflow orchestration tools (e.g., Temporal, Cadence) for long-lived, manifest-driven workflows.
  • Experience building software control planes or orchestration systems with state modeling and reconciliation (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
  • A product mindset with experience building internal platforms or APIs consumed by other engineering teams.
  • Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure.
  • Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE).
  • Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
  • Systems programming in Rust or Go.

Skills

  • Go
  • Python
  • Rust
  • Temporal
  • Cadence
  • Kubernetes controllers/operators
  • Kafka
  • NATS
  • SQS
  • PXE/iPXE
  • Redfish/IPMI
  • BMC
  • VLANs
  • BGP
  • NCCL
  • CUDA
  • InfiniBand/RoCE

Salary/Compensations

  • $240,000 - $280,000

Benefits

  • Startup equity
  • Health insurance
  • Other competitive benefits

About the Company

  • Together AI is a research-driven artificial intelligence company.
  • We believe open and transparent AI systems will drive innovation and create the best outcomes for society.
  • Our mission is to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
  • We have contributed to leading open-source research, models, and datasets to advance the frontier of AI.
  • Our team has been behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama.
  • Join a passionate group of researchers and engineers in building the next generation AI infrastructure.

Equal Opportunity

  • Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.