Lead Site Reliability Engineer (GCP & Hybrid Cloud) Hybrid at Cisco | San Jose, Texas, United States | Rezi

Lead Site Reliability Engineer (GCP & Hybrid Cloud) Hybrid at Cisco

Lead Site Reliability Engineer (GCP & Hybrid Cloud) Hybrid

Cisco · San Jose, Texas, United States

2 months ago

Lead Site Reliability Engineer (GCP & Hybrid Cloud) Hybrid

Cisco · San Jose, Texas, United States

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a Lead SRE, you will own the architectural integrity of our hybrid cloud infrastructure, ensuring our GCP and on-premise Kubernetes environments are resilient and secure. You will set the standard for automation and reliability that enables our AI models to scale globally.

Responsibilities

  • Lead architectural design of scalable hybrid-cloud environments, managing GCP and On-premise Kubernetes clusters with Anthos Service Mesh (ASM) and Istio.
  • Direct implementation of Identity and Access Management (IAM) policies and GCP Quota management for secure and cost-effective resource utilization.
  • Architect multi-region, load-balanced microservices with DDoS hardening, end-to-end encryption, and automated secrets management.
  • Design a comprehensive observability strategy using Elasticsearch and Kibana for proactive alerts on service performance and cost envelope management.
  • Partner with development leads to integrate "Security by Design" into automation and AI agent lifecycle using Apigee for secure API management.

Requirements

  • Expert-level proficiency with Terraform, Kubernetes (GKE & On-prem), and Docker.
  • Hands-on expertise with Anthos Service Mesh (ASM), Istio, and Apigee.
  • Deep understanding of IAM implementation and GCP Quota management.
  • Experience with the ELK stack (Elasticsearch/Kibana) for large-scale observability.
  • Strong financial acumen for cloud cost optimization and proactive budget alerting.
  • Experience managing complex traffic between cloud platforms and on-premise data centers.

Skills

  • Terraform
  • Kubernetes (GKE & On-prem)
  • Docker
  • Anthos Service Mesh (ASM)
  • Istio
  • Apigee
  • IAM implementation
  • GCP Quota management
  • Elasticsearch
  • Kibana
  • Cloud cost optimization
  • Budget alerting
  • Hybrid cloud traffic management

Location

  • hybrid position
  • U.S. and/or Canada locations
  • New York City Metro Area
  • Non-Metro New York state
  • Washington state

Work Type

  • hybrid

Experience Level

  • 7+ years of experience in Cloud/On-prem Operations, SRE, or DevOps.
  • Expert-level proficiency

Education Level

  • Bachelor’s Degree in Computer Science, Engineering, or a related field.
  • GCP Professional Cloud Security Engineer or Network Engineer certification.

Salary/Compensations

  • $165,000.00 to $241,400.00 (U.S. and/or Canada locations)
  • New York City Metro Area: $165,000.00 - $277,600.00
  • Non-Metro New York state & Washington state: $146,700.00 - $247,000.00

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • 401(k) plan with a Cisco matching contribution
  • Paid parental leave
  • Short-term disability coverage
  • Long-term disability coverage
  • Basic life insurance
  • Grants of Cisco restricted stock units
  • 10 paid holidays per full calendar year
  • 1 floating holiday (non-exempt employees)
  • 1 paid day off for employee’s birthday
  • Paid year-end holiday shutdown
  • 4 paid days off for personal wellness
  • 16 days of paid vacation time per full calendar year (non-exempt employees)
  • Flexible vacation time off program (exempt employees)
  • 80 hours of sick time off provided on hire date and each January 1st thereafter
  • Up to 80 hours of unused sick time carried forward
  • Additional paid time away for critical or emergency family issues
  • Optional 10 paid days per full calendar year to volunteer
  • Annual bonuses (non-sales roles)
  • Performance-based incentive pay (sales plans)

About the Company

  • Cisco’s Enterprise AI team enables Generative AI powered experiences across Cisco.
  • The mission is to build secure, scalable AI platforms that empower teams to safely develop, deploy, and operationalize AI-powered solutions.
  • The team operates at the intersection of applied AI, cloud infrastructure, and security.
  • Partners across engineering, security, compliance, and product teams to bring trusted AI to life at an enterprise scale.
  • A fast-growing, highly collaborative team of platform engineers, AI engineers, and data scientists valuing technical depth, ownership, and pragmatic execution.
  • Offers the opportunity to define how secure Generative AI is built and governed inside a global technology leader.
  • Cisco revolutionizes how data and infrastructure connect and protect organizations in the AI era.
  • Innovating fearlessly for 40 years to create solutions that power how humans and technology work together across physical and digital worlds.
  • Solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
  • Fueled by technology depth and breadth, Cisco experiments and creates meaningful solutions.
  • A worldwide network of doers and experts offers limitless opportunities to grow and build.
  • The team collaborates with empathy to make big things happen on a global scale.
  • Cisco's solutions and impact are everywhere.