About the Role
Join the Platform Engineering & AI Operations team, operating at the intersection of Site Reliability Engineering and intelligent infrastructure operations. This role shapes how the bank operates, monitors, and self-heals its cloud platforms, working on real problems at enterprise scale to reduce toil and build automation.
Responsibilities
- Support highly scalable, secure, and highly available architectures across private and public cloud platforms.
- Write code and scripts to automate infrastructure workflows and eliminate toil.
- Extend self-healing automation capabilities built on Ansible, automating routine operational tasks.
- Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points.
- Collaborate with platform teams to provide technical feedback, contribute code changes, and establish data standards and pipelines.
- Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform.
- Minimize risk of reliability failures related to durability, availability, performance, and correctness.
- Participate in on-call rotation for platform support, incident management, and troubleshooting.
Requirements
- 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
- Strong working knowledge of Kubernetes/OpenShift administration and troubleshooting in enterprise environments.
- Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
- Proficiency in Python scripting.
- Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
- Experience with incident management processes, on-call rotations, and post-incident review practices.
- Familiarity with capacity planning, threshold-based alerting, and performance trend analysis.
- Understanding of security and compliance fundamentals.
- Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
- Hands-on experience with public cloud platforms (AWS, Azure, GCP) in hybrid or multi-cloud environments.
- Experience with GPU/compute infrastructure for ML inference workloads.
Skills
- Agile Methodology
- Ansible Tower
- Group Problem Solving
- IT System Administration
- IT Systems Integration
- Kubernetes
- Linux
- Organizational Leadership
- Product Services
- RedHat OpenShift Administration
- Red Hat OS Administration
- Software Development Life Cycle (SDLC)
- System Applications
- System Integration Testing (SIT)
- Systems Software
Location
- Toronto, Canada
Work Type
- Full time
Experience Level
- 5+ years
Benefits
- A comprehensive Total Rewards Program including bonuses and flexible benefits
- Competitive compensation
- Commissions
- Stock where applicable
- Leaders who support your development through coaching and managing opportunities
- Ability to make a difference and lasting impact
- Work in a dynamic, collaborative, progressive, and high-performing team
- Flexible work/life balance options
- Opportunities to do challenging work
- Opportunities to take on progressively greater accountabilities
- Access to a variety of job opportunities across business
About the Company
- At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC.
- We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world.
- Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities.
- RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.
