About the Role
Manage, scale, and evolve our Google Kubernetes Engine platform, ensuring it is production-ready, secure, and continuously improving through deep Kubernetes expertise and strong cloud delivery practices.
Responsibilities
- Provision, manage, and optimize GKE clusters including Autopilot and Standard modes across multiple environments.
- Design and implement cluster networking, including VPC-native clusters, Shared VPC, and Private GKE configurations.
- Manage GKE IAM, RBAC policies, and Workload Identity integration to enforce least-privilege access.
- Own cluster lifecycle management: upgrades, node pool management, auto-scaling, and cost optimization.
- Define and maintain infrastructure-as-code for GKE using Terraform and Helm charts.
- Implement GitOps workflows using ArgoCD or Flux for continuous delivery to GKE environments.
- Configure and manage Kubernetes-native tooling: Ingress controllers, service mesh (Istio/Anthos), and policy controllers (OPA/Gatekeeper).
- Deliver observability solutions using Google Cloud Monitoring, Managed Prometheus, and Cloud Logging.
- Automate operational tasks and cluster configuration using Ansible playbooks.
- Collaborate with application teams to define workload specifications, resource quotas, and namespace governance.
- Support cloud delivery activities including release planning, environment management, and go-live support.
Requirements
- 4+ years of experience administering production Kubernetes clusters, with at least 2 years on GKE.
- Strong understanding of GKE networking: VPC-native, Shared VPC, Private Clusters, Cloud NAT.
- Hands-on experience with Terraform for GKE cluster provisioning and node pool management.
- Proficiency in Helm for Kubernetes application packaging and deployment.
- Experience with GKE IAM, RBAC, and Workload Identity Federation.
- Familiarity with GitOps tools (ArgoCD, Flux) and CI/CD pipeline integration.
- Working knowledge of Ansible for configuration management and automation.
- Strong understanding of container security, image scanning, and Binary Authorization.
- Google Professional Cloud DevOps Engineer or CKA/CKAD certification.
- Experience with Anthos Service Mesh or Istio on GKE.
- Familiarity with Vertex AI Workbench or ML workloads on GKE.
- Experience managing multi-tenant GKE platforms at enterprise scale.
- Background in SRE practices, SLO definition, and error budget management.
Skills
- Kubernetes
- Google Kubernetes Engine (GKE)
- Cloud Delivery
- Terraform
- Helm
- ArgoCD
- Flux
- Ansible
- Istio
- Anthos
- OPA/Gatekeeper
- Google Cloud Monitoring
- Managed Prometheus
- Cloud Logging
- VPC-native networking
- Shared VPC
- Private GKE configurations
- IAM
- RBAC
- Workload Identity
- GitOps
- CI/CD
- Configuration Management
- Container Security
- Image Scanning
- Binary Authorization
- SRE Practices
- SLO Definition
- Error Budget Management
Experience Level
- 4+ years of experience administering production Kubernetes clusters
- at least 2 years on GKE
