About the Role
We are seeking a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. This role is crucial for ensuring the reliability, scalability, and performance of our infrastructure, with a strong emphasis on Google Cloud Platform (GCP). You will be responsible for architecting and maintaining high-availability systems, automating operational tasks, and ensuring our services can handle millions of voice AI interactions.
Responsibilities
- Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
- Architect and automate CI/CD pipelines for rapid, reliable deployments.
- Implement robust monitoring, alerting, and observability strategies.
- Partner with engineering teams to optimize performance, cost, and reliability of backend services.
- Drive incident response, post-mortem analysis, and long-term remediation efforts.
- Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
- Collaborate with cross-functional teams on infrastructure roadmaps and security standards.
- Lead department-wide compliance (PCI, SOC) initiatives.
Requirements
- 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
- Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
- Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
- Deep experience with Kubernetes, container orchestration, and service mesh architectures.
- Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
- Experience designing and managing high-throughput, distributed systems.
- Strong problem-solving skills and a growth mindset.
- Excellent communication skills and ability to mentor engineers.
Skills
- Google Cloud Platform (GCP)
- Infrastructure as Code (IaC)
- Terraform
- Pulumi
- Kubernetes
- Container orchestration
- Service mesh architectures
- Monitoring tools
- Observability tools
- Datadog
- Prometheus
- Grafana
- Cloud Monitoring
- Distributed systems
- Clojure/ClojureScript
Location
- Canada
Work Type
- Remote
Experience Level
- Staff
- 12+ years
Salary/Compensations
- Salary range provided based on location and experience
Benefits
- Equity
- Comprehensive healthcare
- Paid time off
About the Company
- We are a company focused on Retail and Restaurants AI.
