Member of Technical Staff, Site Reliablity Engineer at Vapi | San Francisco, CA | Rezi

Member of Technical Staff, Site Reliablity Engineer at Vapi

Member of Technical Staff, Site Reliablity Engineer

Vapi · San Francisco, CA

2 months ago

Member of Technical Staff, Site Reliablity Engineer

Vapi · San Francisco, CA

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Vapi is seeking an individual to drive 99.99% call completion by managing incident command, owning SLOs and error budgets, and fostering a reliability culture. This role involves shipping code for services that monitor and manage the platform, including auto-remediation, capacity forecasters, and oncall tooling, as well as capacity planning, load testing, and autoscaling.

Responsibilities

  • Join the oncall rotation within 30 days.
  • Analyze past incidents to create a prioritized reliability backlog.
  • Define the initial set of SLOs for the call-completion path.
  • Establish error budgets and SLO-based alerting in Chronosphere/Prometheus for key services within 60 days.
  • Conduct the first comprehensive load test against provider rate limits and per-org concurrency.
  • Tune autoscaling for wscaler and workerpool-cron-scaler.
  • Ship a platform service (capacity forecaster, auto-remediation, or oncall tooling) in Go or TypeScript within 90 days.
  • Own the postmortem process.
  • Drive measurable improvements in p99 call completion or MTTR.

Requirements

  • Experience running incident command and postmortem discipline at scale on an oncall rotation.
  • Experience operating SLOs and error budgets using tools like Chronosphere, Prometheus, Grafana, or Datadog.
  • Experience with capacity planning and load testing for production systems with live users.
  • Proficiency in Kubernetes production operations, including pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, and graceful shutdown.
  • Knowledge of backpressure and autoscaling patterns such as KEDA and custom metrics scaling.
  • Ability to build platform services in Go or TypeScript.
  • Experience in real-time/latency-sensitive product environments where performance degradation leads to dropped calls.

Skills

  • Go
  • TypeScript
  • Bash
  • Chronosphere
  • Prometheus
  • Grafana
  • Datadog
  • OpenTelemetry
  • Kubernetes
  • EKS
  • HPA/VPA tuning
  • PodDisruptionBudgets
  • Graceful shutdown
  • Pod crash diagnosis
  • KEDA
  • Custom metrics scaling
  • Load testing
  • Capacity planning

Location

  • Remote

Work Type

  • Full-time

Experience Level

  • Mid-level
  • Senior

Salary/Compensations

  • Competitive salary

Benefits

  • Excellent equity ownership
  • Comprehensive health coverage (medical, dental, vision)
  • Quarterly off-sites
  • Flexible time off
  • Catered meals
  • Transportation
  • Gym membership
  • $10k annual L&D budget

About the Company

  • Vapi provides voice AI that resolves customer issues without transfers, enabling voice agents that understand businesses and resolve problems efficiently.
  • The company has experienced significant growth, serving customers like Amazon Ring, ServiceTitan, and Intuit.
  • Vapi recently raised $50M in Series B funding, bringing the total raised to $72M, with support from prominent investors.
  • Vapi offers a generational impact opportunity to build the human interface for businesses.
  • The company fosters an ownership culture with 70% of the team being previous founders.
  • Vapi is backed by Tier-1 Investors including YC, KP seed, and Bessemer.