Staff Site Reliability Engineer at SoundHound AI | Toronto, Canada | Rezi

Staff Site Reliability Engineer at SoundHound AI

Staff Site Reliability Engineer

SoundHound AI · Toronto, Canada

1 weeks ago

Staff Site Reliability Engineer

SoundHound AI · Toronto, Canada

12 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. This role is crucial for ensuring the reliability, scalability, and performance of our infrastructure, with a strong emphasis on Google Cloud Platform (GCP). You will be responsible for architecting and maintaining high-availability systems, automating operational tasks, and ensuring our services can handle the demands of millions of voice AI interactions.

Responsibilities

  • Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • Architect and automate CI/CD pipelines for rapid, reliable deployments.
  • Implement robust monitoring, alerting, and observability strategies.
  • Partner with engineering teams to optimize performance, cost, and reliability of backend services.
  • Drive incident response, post-mortem analysis, and long-term remediation efforts.
  • Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
  • Collaborate with cross-functional teams on infrastructure roadmaps and security standards.
  • Lead department-wide compliance (PCI, SOC) initiatives.

Requirements

  • 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
  • Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Deep experience with Kubernetes, container orchestration, and service mesh architectures.
  • Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
  • Experience designing and managing high-throughput, distributed systems.
  • Strong problem-solving skills and a growth mindset.
  • Excellent communication skills and ability to mentor engineers.
  • Experience working in a high-velocity, customer-focused environment.
  • Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
  • Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
  • Experience implementing security and compliance best practices in the cloud.

Skills

  • Google Cloud Platform (GCP)
  • GKE
  • Compute Engine
  • Cloud Run
  • Pub/Sub
  • Terraform
  • Pulumi
  • Kubernetes
  • Container Orchestration
  • Service Mesh
  • Datadog
  • Prometheus
  • Grafana
  • Cloud Monitoring
  • Distributed Systems
  • Problem-solving
  • Communication
  • Mentoring
  • Functional Programming
  • Clojure
  • ClojureScript
  • Cloud Security
  • Cloud Compliance

Location

  • Canada

Work Type

  • Remote

Experience Level

  • Staff
  • 12+ years

Salary/Compensations

  • Salary range based on location and years of experience

Benefits

  • Equity
  • Comprehensive healthcare
  • Paid time off

About the Company

  • Our Retail and Restaurants AI team focuses on leveraging AI for voice interactions.