About the Role
The Site Reliability Engineering (SRE) team is central to delivering seamless and robust services. This role involves solving complex challenges, driving innovation, and improving observability while reducing toil.
Responsibilities
- Analyze and forecast system capacity requirements for scalability and performance during high-profile events.
- Participate in incident response, conduct post-incident reviews, and implement improvements to monitoring.
- Develop and maintain monitors and alerts across all services.
- Optimize system performance through tuning and configuration adjustments.
- Develop and maintain disaster recovery plans and procedures for business continuity.
- Create and maintain comprehensive documentation for systems, processes, and procedures.
- Monitor and optimize cloud infrastructure for efficient resource utilization.
- Identify opportunities for automation and implement solutions to reduce manual intervention.
- Collaborate with cross-functional teams to align on goals and deliver high-quality solutions.
- Act as a technical lead, directing workflow as needed.
- Exercise independent judgment and discretion in significant matters.
- Maintain consistent and punctual attendance.
- Perform other assigned duties and responsibilities.
Requirements
- Demonstrate understanding of wider operational performance factors influenced by underlying infrastructure workload (server platforms, databases, networking).
- Exhibit a 'detective' mindset for diagnosing issues and a passion for detail and investigation.
- Proactively diagnose problems using a holistic knowledge-set and implement permanent fixes through coding, process rewriting, or third-party collaboration.
- View automation as an opportunity to overcome scale challenges.
- Possess a flexible approach to technologies.
Skills
- Cloud platforms (AWS, GCP, Azure)
- Scripting languages (Python, Bash)
- Infrastructure-as-code tools (Terraform, Ansible)
- Containerization and orchestration tools (Docker, Kubernetes)
- Monitoring tools (Datadog, Splunk)
- Database performance monitoring and tuning (NoSQL, SQL)
- Kubernetes performance monitoring and tuning
- Excellent problem-solving skills
- Attention to detail
- Effective communication skills
- Collaboration skills
- CI/CD pipeline management
- Cloud Cost Optimization
- Automating deployment processes
- Front-end development (React)
Location
- Denver, Colorado, USA
Work Type
- Full-time
Experience Level
- 5-7 Years
Education Level
- Bachelor's Degree
Benefits
- Options, expert guidance, and tools to support physical, financial, and emotional well-being.
About the Company
- Comcast brings together the best in media and technology, driving innovation for entertainment and online experiences.
- Comcast Technology Solutions is a software technology company enabling streaming services, TV stations, pay TV operators, content providers, broadband media sites, and mobile businesses.
- Our Cloud Video Platform (CVP) offers a product catalogue that allows customers to securely manage digital media, publish content, and monetize distribution.
- Our customers include Deutsche Telekom, Viaplay, Fox, Disney, NBC, Paramount+, and many others.
Equal Opportunity
- Comcast is an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.
