Senior Site Reliability Engineer -AI Infrastructure Operations at Nscale | San Francisco, CA, US | Rezi

Senior Site Reliability Engineer -AI Infrastructure Operations at Nscale

Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale · San Francisco, CA, US

2 weeks ago

Senior Site Reliability Engineer -AI Infrastructure Operations

Nscale · San Francisco, CA, US

20 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

This is a senior SRE role focused on owning and improving the reliability of critical production services. The goal is to reduce the frequency of incidents through automation, robust design, and mentoring other engineers.

Responsibilities

  • Own reliability for critical production services end to end; set the direction, not just respond to what breaks.
  • Grow the team, not just the systems; mentor other SREs through design review, pairing, and incident debriefs, and hold the bar that pulls everyone up to it.
  • Set the standards the rest of the team works to: the SLO framework, the incident process, and the on-call practices that keep it sustainable.
  • Get in early on design reviews and architecture decisions, so reliability is built in rather than bolted on after the first outage.
  • Lead the hardest incidents and the root causes nobody else can crack; turn each one into a change that keeps it from coming back.
  • Build the tooling and automation that removes toil for the whole team, not just your own surface area.

Requirements

  • 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production at scale in a data center or cloud environment.
  • Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt, not scripts that run once and rot.
  • Deep command of Linux, networking, and distributed systems, plus the judgment to know where the real failure modes hide.
  • Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the metal, not just the cloud console.
  • Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to get there fast.
  • Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale, incident process, on-call that people can actually live with.
  • A track record as the senior voice in incidents and design reviews, trusted to make the call under pressure.
  • A habit of raising the people around you without being asked to.

Skills

  • Python
  • Go
  • Linux
  • Networking
  • Distributed systems
  • Kubernetes
  • Virtualized environments
  • Bare-metal environments
  • AI workloads
  • GPU workloads
  • High-performance computing (HPC)
  • SLOs
  • Observability
  • Alerting
  • Incident process
  • On-call practices
  • High-performance networking (InfiniBand, RDMA)

Experience Level

  • Senior

Salary/Compensations

  • $170,000 - $265,000 USD

Benefits

  • Competitive base plus equity, reviewed every 12 months.
  • Real ownership from the start, and a direct hand in how reliability works across the platform.
  • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape your day.
  • Medical
  • Dental
  • Vision
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation

About the Company

  • Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work.

Equal Opportunity

  • At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
  • If there’s anything we can do to accommodate your specific situation, please let us know.