About the Role
As an Incident Commander on the site reliability team, you will work cross-functionally across engineering, acting as the front line for incidents and collaborating with Release Engineering to prevent future events. This role is responsible for managing all incidents across online and physical organizations, including classification, documentation, investigation, escalation, diagnosis, recovery, and root cause analysis. You will also drive improvements to service delivery and release processes based on disruption reports.
Responsibilities
- Drive and enhance collaboration with other Incident Commanders, Customer Support, Application and Engineering teams for real-time incident management.
- Provide leadership for developing practices, frameworks, process flows, templates, and process guides.
- Continuously improve and enhance the internal framework, methodology, processes, and tools.
- Develop and maintain key practical capabilities.
- Collaborate with SRE Teams and Infrastructure teams to identify requirements and gaps resulting in downtime or blindspots.
- Recommend innovative solutions that enable the organization to deliver on its objectives and goals.
- Promote opportunities for Continuous Service Improvements.
- Manage and update Root Cause Analysis documentation.
- Lead SRE communications to stakeholders via E-mail, Slack, & Teams in a timely manner.
- Lead initiatives to promote JIRA Release Ticket management, quality, and alignment with Incident management communication supporting SLAs.
- Perform other duties as required.
Requirements
- Experience in a similar role or incident management role.
- Experience and understanding of Containerization (Docker & Kubernetes preferred).
- Understanding of automation tools such as configuration management and infrastructure as code tools (Terraform, Ansible, Helm, etc.).
- Experience with a programming language.
- Comfortable working within Linux environments.
- Experience working with AWS, GCP, and on-premises environments.
- Ability to work independently and learn quickly with little supervision.
- Ability to handle multiple projects simultaneously.
- Willingness to drop everything and take on an ad-hoc task.
- Tech-savvy and passionate about learning new technologies and tools.
- Outgoing, and able to keep a conversation going naturally to extract needed information.
Skills
- Docker
- Kubernetes
- Terraform
- Ansible
- Helm
- AWS
- GCP
- Linux
- Postgres
- MySQL
- Elastic Search
- Kafka
- Redis
- Terragrunt
- Prometheus
- Python
- Talos Linux
Location
- Remote
Work Type
- Full-time
Experience Level
- Mid-level
Education Level
- Degree in computer science, engineering, and/or similar experience.
Salary/Compensations
- $90,000—$135,000 CAD
Benefits
- Competitive compensation package.
- Fun, relaxed work environment.
- Education and conference reimbursements.
- Parental leave top up.
- Opportunities for career progression and mentoring others.
- Best-in-class benefits designed to support employees physically, financially, and emotionally.
About the Company
- PENN Entertainment, Inc. is North America’s leading provider of integrated entertainment, sports content, and casino gaming experiences.
- We deliver experiences through casinos, racetracks, online gaming, sports betting, and entertainment content.
- Our cutting-edge online gaming and sports media products are powered by proprietary in-house technology.
- We are committed to supporting employee career growth, skill expansion, and exploration of new opportunities.
- We have locations throughout North America.
Equal Opportunity
- Penn Interactive is proud to be an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.
