About the Role
The Database Reliability Engineering team builds and operates production database and data systems for Cisco Meraki's cloud platform. This role focuses on making high-volume data ingestion, customer data access, and operational data needs reliable at a global scale, impacting engineers and customers.
Responsibilities
- Plan, administer, maintain, and secure PostgreSQL infrastructure in collaboration with Site Reliability Engineering (SRE) teams.
- Design, build, and maintain ETL pipelines for PostgreSQL, and develop procedures and scripts for data migration.
- Perform operational database administration tasks including installation, upgrades, patching, backup/recovery, monitoring, capacity planning, and architectural changes in cloud environments.
- Participate in production operations such as on-call rotation, incident response, monitoring, alerting, and post-incident review processes.
Requirements
- 5+ years of experience designing, operating, and troubleshooting PostgreSQL in production environments.
- 5+ years of experience managing production database or distributed data systems across application, database, operating system, storage, and network layers.
- 3+ years of experience in Linux systems engineering (performance tuning, memory management, I/O tuning, configuration, security, and networking).
- 3+ years of experience automating infrastructure or database operations with tools such as Terraform, Ansible, Chef, or Puppet.
- 2+ years of experience using at least one scripting or programming language such as Python, Bash, Go, Ruby, or Perl for automation and operational tooling.
- Experience working in a polyglot production data environment, including at least one non-PostgreSQL system such as Kafka/MSK, ClickHouse, Redis, MySQL, Cassandra, Elasticsearch, or a similar distributed data system.
- Experience building database platform tooling, self-service workflows, and paved paths that enable application teams to use data systems safely and efficiently.
- Previous work with cloud infrastructure and managed data services, including PostgreSQL on Kubernetes, Amazon RDS, AWS, Terraform, service discovery, and secrets management.
- Ability to operate MSK/Kafka, ClickHouse, or Redis at scale, covering cluster operations, replication, partitioning or sharding, retention, capacity planning, and workload tuning.
- Previous responsibility for refining reliability practices for production data systems, such as SLOs, disaster recovery plans, backup validation, and failover testing.
- Deep knowledge of PostgreSQL internals and operational behavior, including concurrency, transaction consistency, replication, maintenance, backup and recovery, indexing, and query performance.
- Ability to communicate effectively in writing and verbally by producing design documents, leading operational reviews, explaining tradeoffs, and mentoring engineers.
Skills
- PostgreSQL
- ETL pipelines
- Data migration
- Linux systems engineering
- Terraform
- Ansible
- Chef
- Puppet
- Python
- Bash
- Go
- Ruby
- Perl
- Kafka/MSK
- ClickHouse
- Redis
- MySQL
- Cassandra
- Elasticsearch
- Kubernetes
- Amazon RDS
- AWS
- Service discovery
- Secrets management
- SLOs
- Disaster recovery
- Backup validation
- Failover testing
Location
- U.S.
- Canada
Work Type
- Full-time
Experience Level
- Senior
Salary/Compensations
- $165,000.00 - $241,400.00 (U.S. and/or Canada)
- $165,000.00 - $277,600.00 (New York City Metro Area)
- $146,700.00 - $247,000.00 (Non-Metro New York state & Washington state)
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- 401(k) plan with Cisco matching contribution
- Paid parental leave
- Short-term disability coverage
- Long-term disability coverage
- Basic life insurance
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- 1 paid day off for employee’s birthday
- Paid year-end holiday shutdown
- 4 paid days off for personal wellness
- 16 days of paid vacation time per full calendar year (non-exempt employees)
- Flexible vacation time off program (exempt employees)
- 80 hours of sick time off provided on hire date and each January 1st thereafter
- Up to 80 hours of unused sick time carried forward
- Additional paid time away for critical or emergency family issues
- Optional 10 paid days per full calendar year to volunteer
- Annual bonuses (non-sales roles)
- Performance-based incentive pay (sales roles)
- Cisco restricted stock units
About the Company
- At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond.
- We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds.
- Our solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
- Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions.
- We work as a team, collaborating with empathy to make really big things happen on a global scale.
- Our solutions are everywhere, so our impact is everywhere.
- We are Cisco, and our power starts with you.
