Staff Site Reliability Engineer (Copy) at SoundHound AI | CA | Rezi

Staff Site Reliability Engineer (Copy) at SoundHound AI

Staff Site Reliability Engineer (Copy)

SoundHound AI · CA

3 weeks ago

Staff Site Reliability Engineer (Copy)

SoundHound AI · CA

23 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. This role is crucial for ensuring the reliability, scalability, and performance of our infrastructure, with a strong emphasis on Google Cloud Platform (GCP). You will be responsible for architecting and maintaining high-availability systems, automating operational tasks, and ensuring our services can handle millions of voice AI interactions.

Responsibilities

  • Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • Architect and automate CI/CD pipelines for rapid, reliable deployments.
  • Implement robust monitoring, alerting, and observability strategies.
  • Partner with engineering teams to optimize performance, cost, and reliability of backend services.
  • Drive incident response, post-mortem analysis, and long-term remediation efforts.
  • Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
  • Collaborate with cross-functional teams on infrastructure roadmaps and security standards.
  • Lead department-wide compliance (PCI, SOC) initiatives.

Requirements

  • 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
  • Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Deep experience with Kubernetes, container orchestration, and service mesh architectures.
  • Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
  • Experience designing and managing high-throughput, distributed systems.
  • Strong problem-solving skills and a growth mindset.
  • Excellent communication skills and ability to mentor engineers.

Skills

  • Google Cloud Platform (GCP)
  • Infrastructure as Code (IaC)
  • Terraform
  • Pulumi
  • Kubernetes
  • Container orchestration
  • Service mesh architectures
  • Monitoring tools
  • Observability tools
  • Datadog
  • Prometheus
  • Grafana
  • Cloud Monitoring
  • Distributed systems
  • Clojure/ClojureScript

Location

  • Canada

Work Type

  • Remote

Experience Level

  • Staff
  • 12+ years

Salary/Compensations

  • Salary range provided based on location and experience

Benefits

  • Equity
  • Comprehensive healthcare
  • Paid time off

About the Company

  • We are a company focused on Retail and Restaurants AI.