Senior Backend Engineer, Managed Kubernetes & Slurm at Lightning AI | CA, US | Rezi

Senior Backend Engineer, Managed Kubernetes & Slurm at Lightning AI

Senior Backend Engineer, Managed Kubernetes & Slurm

Lightning AI · CA, US

5 days ago

Senior Backend Engineer, Managed Kubernetes & Slurm

Lightning AI · CA, US

5 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a Senior Backend Engineer on the Managed Services team, you'll build the control planes, backend services, and automation that provision, operate, and scale Kubernetes and Slurm clusters across our GPU fleet. Your work will turn raw GPU capacity into production-ready environments where customers can seamlessly train models, run inference, and deploy AI workloads. This role involves solving challenging distributed systems and cloud infrastructure problems, building automation and platform capabilities for scalable, observable, and resilient Kubernetes and Slurm environments.

Responsibilities

  • Design, build, and operate backend services in Go or Python that power Lightning AI's managed infrastructure platform.
  • Develop control plane services that provision, orchestrate, and manage Kubernetes and Slurm clusters across large-scale GPU infrastructure.
  • Build distributed systems that automate cluster lifecycle management, workload scheduling, infrastructure provisioning, and platform operations.
  • Develop platform capabilities using Kubernetes APIs, controllers, operators, and other cloud-native technologies.
  • Improve the reliability, scalability, security, and observability of our managed platform through automation and operational excellence.
  • Diagnose and resolve complex production issues across Kubernetes, distributed systems, networking, and cloud infrastructure.
  • Collaborate with infrastructure, AI, and platform engineering teams to shape the future of our cloud platform.
  • Contribute to technical design, architecture, mentoring, engineering best practices, and on-call operations.

Requirements

  • Significant professional experience designing, building, and operating production backend systems using Go or Python.
  • Deep hands-on experience with Kubernetes or Slurm, including operating large-scale production environments.
  • Strong understanding of distributed systems, cloud-native architectures, and production infrastructure.
  • Experience designing and building scalable backend services, APIs, and automation for infrastructure or platform operations.
  • Strong understanding of networking, storage, and cloud infrastructure fundamentals.
  • Familiarity with observability, CI/CD, testing, production operations, and incident response.
  • Ability to own complex technical projects while collaborating effectively across engineering teams.
  • Experience with Kubernetes platform development, including operators, controllers, CRDs, or control plane components.
  • Experience with Slurm administration, scheduling, or HPC environments.
  • Experience with Infrastructure-as-Code or GitOps tooling such as Terraform, Crossplane, Helm, Argo CD, Flux, or Kustomize.
  • Experience with GPU infrastructure, AI platforms, or large-scale compute environments.
  • Experience with Kubernetes networking, storage, or multi-tenant platform architecture.
  • Experience with event-driven systems, gRPC, or distributed messaging.
  • Experience building CI/CD systems for production environments.
  • Contributions to cloud-native or open source infrastructure projects.

Skills

  • Go
  • Python
  • Kubernetes
  • Slurm
  • Distributed systems
  • Cloud-native architectures
  • Backend services
  • APIs
  • Automation
  • Networking
  • Storage
  • Cloud infrastructure
  • Observability
  • CI/CD
  • Testing
  • Production operations
  • Incident response
  • Kubernetes APIs
  • Controllers
  • Operators
  • CRDs
  • Control plane components
  • HPC environments
  • Infrastructure-as-Code
  • GitOps
  • Terraform
  • Crossplane
  • Helm
  • Argo CD
  • Flux
  • Kustomize
  • GPU infrastructure
  • AI platforms
  • Large-scale compute environments
  • Event-driven systems
  • gRPC
  • Distributed messaging

Location

  • San Francisco
  • New York City
  • Seattle

Work Type

  • Hybrid

Experience Level

  • Senior

Salary/Compensations

  • $180,000—$250,000 USD

Benefits

  • Discretionary bonus
  • Meaningful equity
  • Comprehensive health coverage (medical, dental, vision)
  • RSUs
  • 401(k) matching (U.S.)
  • Pension contributions (U.K.)
  • Unlimited PTO
  • Company holidays
  • Floating holidays
  • Company-wide winter break
  • Paid parental & family leave
  • Annual learning and development allowance
  • Wellness and work-from-home stipends
  • Four weeks of paid sabbatical leave after four years of service
  • Flexible schedules
  • Complimentary meals at office hubs

About the Company

  • Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems.
  • Through its merger with Voltage Park, Lightning AI combines developer-first software with cost-efficient, large-scale compute.
  • The company serves solo researchers, startups, and large enterprises globally.
  • Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London.
  • The company is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
  • Lightning AI runs one of the largest AI-native clouds, with 25,000+ H100, B200, and GB300 GPUs.

Equal Opportunity

  • Lightning AI is committed to fostering an inclusive and diverse workplace.
  • Diverse teams drive innovation and create better products.
  • Equal employment opportunities are provided to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic.
  • The company is dedicated to building a culture where everyone can thrive and contribute to their fullest potential.