About the Role
Varo’s SRE team designs, builds, and runs large-scale, distributed, fault-tolerant systems powering Varo's operations. The team focuses on automation and observability, striving to automate manual tasks and promote a data-driven approach to platform scaling. Daily activities include scaling production infrastructure, building CI/CD pipelines, and collaborating with developers to enhance the platform.
Responsibilities
- Manage, upgrade, and autoscale EKS clusters across multiple environments (SIT, UAT, Prod) and AWS accounts.
- Write Terraform modules and Helm charts to support GitOps workflows using ArgoCD and GitLab CI/CD pipelines.
- Maintain and troubleshoot Kafka (MSK) clusters, including broker health, connectors, and CDC pipelines.
- Improve observability using Prometheus, Thanos, Grafana, and ELK while proactively identifying cloud cost-optimization opportunities.
- Automate operational tasks with Python and leverage AI/ML techniques for predictive alerting and intelligent runbooks.
- Handle Platform Service Desk requests, including Terraform merge request reviews, access management, and deployment support.
- Participate in the production on-call rotation, support incident response, and contribute to blameless post-mortems.
Requirements
- 3+ years of experience in an SRE, DevOps, or Infrastructure Engineering role.
- Ability to work independently and manage multiple workstreams.
- Strong hands-on experience with core AWS services, including EKS, EC2, RDS Aurora, MSK, S3, IAM, VPC, and Direct Connect.
- Deep production experience with Kubernetes (upgrades, networking, RBAC) alongside Helm and GitOps tools like ArgoCD.
- Advanced proficiency with Terraform, including writing modules and managing multi-account/multi-environment states.
- Experience supporting and maintaining data platforms such as Airflow, Databricks, EMR, Kafka/MSK, or CDC pipelines.
- Solid understanding of networking (VPCs, security groups, Istio, DNS).
- Strong Python scripting skills for tooling and automation.
- Experience managing observability stacks (Prometheus, Grafana, ELK).
- Effectively leverage AI/LLM tools for automation and incident analysis.
- Experience with Karpenter and KEDA (Nice to Have).
- GitLab CI/CD pipeline experience (Nice to Have).
- Hashicorp Vault for secrets management (Nice to Have).
Skills
- AWS
- EKS
- EC2
- RDS Aurora
- MSK
- S3
- IAM
- VPC
- Direct Connect
- Kubernetes
- Helm
- ArgoCD
- Terraform
- GitOps
- Python
- Prometheus
- Thanos
- Grafana
- ELK
- Kafka
- CDC pipelines
- Istio
- DNS
- AI/ML
- Karpenter
- KEDA
- GitLab CI/CD
- Hashicorp Vault
Experience Level
- 3+ years
About the Company
- Varo is an entirely new kind of bank. All digital, mission-driven, FDIC insured and designed for the way our customers live their lives. A bank for all of us.
Equal Opportunity
- Varo is an equal opportunity employer. Varo embraces diversity and we are committed to building teams that represent a variety of backgrounds, perspectives, and skills. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.
