About the Role
We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.
Responsibilities
- Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
- Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
- Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
- Participate in and improve our 24x7 incident response and on-call rotation.
- Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
- Automate production operations to provide guardrails and continuous platform operation.
- Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
- Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
- Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
- Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
- Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
- Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.
Requirements
- 5+ years of experience in a related role
- Proficiency in software development with languages such as Python or Go
- Shown ability to build and implement scalable, well-tested, and security-focused solutions that integrate security protocols throughout the development and deployment lifecycle
- Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols
- Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs
Skills
- Kubernetes
- Service Mesh
- Prometheus
- OpenTelemetry
- ArgoCD
- AWS
- Python
- Go
- Unix/Linux systems
Location
- San Francisco
- Seattle
- Austin
- New York
Work Type
- Hybrid
Experience Level
- Senior
- 5+ years
Salary/Compensations
- $165,000.00 - $241,400.00 (U.S. and/or Canada)
- $165,000.00 - $277,600.00 (New York City Metro Area)
- $146,700.00 - $247,000.00 (Non-Metro New York state & Washington state)
Benefits
- Medical insurance
- Dental insurance
- Vision insurance
- 401(k) plan with Cisco matching contribution
- Paid parental leave
- Short-term disability coverage
- Long-term disability coverage
- Basic life insurance
- Restricted stock units
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- 1 paid day off for employee’s birthday
- Paid year-end holiday shutdown
- 4 paid days off for personal wellness
- 16 days of paid vacation time per full calendar year (non-exempt)
- Flexible vacation time off program (exempt)
- 80 hours of sick time off provided on hire date and each January 1st thereafter
- Up to 80 hours of unused sick time carried forward
- Additional paid time away for critical or emergency family issues
- Optional 10 paid days per full calendar year to volunteer
- Annual bonuses (non-sales roles)
- Performance-based incentive pay (sales roles)
About the Company
- Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership.
- Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences.
- ThousandEyes is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios.
- At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond.
- We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds.
- These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
- Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions.
- Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless.
- We work as a team, collaborating with empathy to make really big things happen on a global scale.
- Because our solutions are everywhere, our impact is everywhere.
- We are Cisco, and our power starts with you.
