Senior Site Reliability Engineer - Storage at Lambda | CA, US | Rezi

Senior Site Reliability Engineer - Storage at Lambda

Senior Site Reliability Engineer - Storage

Lambda · CA, US

1 months ago

Senior Site Reliability Engineer - Storage

Lambda · CA, US

a month ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Senior Site Reliability Engineer - Storage role.

Rezi rewrites your resume against Lambda's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Senior Site Reliability Engineer - Storage posting at Lambda — free, in seconds.

About the Role

Lambda's Storage Engineering team operates at a massive scale, providing the backbone for world-class storage offerings. We are seeking engineers to solve complex problems, own critical systems, and shape the future of AI infrastructure.

Responsibilities

  • Own the reliability, performance, and capacity health of Lambda's production storage fleet across all data centers.
  • Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
  • Investigate and resolve storage-related incidents using deep telemetry, logs, and performance profiling.
  • Automate ticketing, escalation, and incident-response workflows.
  • Design and maintain self-healing automation for common failure modes.
  • Implement CI/CD pipelines for storage automation and tooling.
  • Partner with Storage Engineers, Fleet Orchestration, and Release Engineering to automate the deployment and configuration of software-defined storage.
  • Work with hardware and networking teams to diagnose low-level I/O and network issues.
  • Participate in an on-call rotation supporting Lambda's storage fleet, focusing on driving down MTTR and building automation.

Requirements

  • 5+ years of experience operating Linux systems in production or HPC environments.
  • Hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS, or similar).
  • Hands-on experience operating Software-Defined Storage (SDS) platforms at scale, including integrating with their management and data-plane APIs.
  • Strong incident-response instincts: comfortable owning a production storage incident end to end.
  • Working experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic — including building dashboards and alert/pager routing.
  • Working experience with Kubernetes (GitOps tooling such as ArgoCD, Helm/Kustomize) and hands-on troubleshooting.
  • Working experience with CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and systems programming in Python or Go.
  • Working experience with Infrastructure as Code (Terraform, Ansible).
  • Solid understanding of core storage protocols across file (NFS, SMB), object (S3), block (NVMe-oF/TCP), and structured (vector DB, SQL) storage.

Skills

  • CEPH
  • Lustre
  • GPFS
  • Software-Defined Storage (SDS)
  • Prometheus
  • Grafana
  • Alertmanager
  • Datadog
  • SumoLogic
  • Kubernetes
  • ArgoCD
  • Helm
  • Kustomize
  • GitHub Actions
  • Jenkins
  • BuildKite
  • Docker
  • Podman
  • Python
  • Go
  • Terraform
  • Ansible
  • NFS
  • SMB
  • S3
  • NVMe-oF/TCP
  • VAST
  • Weka
  • NetApp
  • Dell PowerScale
  • Kubernetes CSI drivers
  • SR-IOV
  • KVM
  • QEMU
  • GPUDirect Storage
  • RDMA
  • InfiniBand
  • RoCE
  • ethtool
  • mlxlink
  • clush

Location

  • San Francisco
  • San Jose

Work Type

  • Hybrid

Experience Level

  • 5+ years of experience operating Linux systems

Salary/Compensations

  • The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

Benefits

  • Generous cash & equity compensation
  • Health, dental, and vision coverage for you and your dependents
  • Wellness and commuter stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible paid time off plan that we all actually use

About the Company

  • Founded in 2012, with 500+ employees, and growing fast.
  • Investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove.
  • Have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG.
  • Values are publicly available: https://lambda.ai/careers.

Equal Opportunity

  • Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.