Site Reliability Engineer Leader at Kyndryl | Ontario | Rezi

Site Reliability Engineer Leader at Kyndryl

Site Reliability Engineer Leader

Kyndryl · Ontario

3 weeks ago

Site Reliability Engineer Leader

Kyndryl · Ontario

22 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

As a Site Reliability Engineer (SRE) at Kyndryl, you will ensure the reliability, resiliency, and innovation of information systems and ecosystems. You will drive continuous improvement, deliver exceptional service, analyze business needs, and provide strategic advice and designs throughout the software lifecycle. Your role involves building trusted relationships with customers and partnering for success, working on end-to-end services across customer sites and platforms. You will embrace an entrepreneurial mindset, focus on quality, robustness, and security, and implement cutting-edge tools to enhance operations and gather feedback. Identifying and mitigating operational issues will be crucial for seamless customer experiences.

Responsibilities

  • Ensure reliability, resiliency, and innovation in information systems and ecosystems.
  • Drive continuous improvement and deliver exceptional service to customers.
  • Analyze business needs, tackle complex problems, and provide strategic advice and designs.
  • Be involved in all stages of the software lifecycle, from building and testing to deploying changes and maintaining systems.
  • Build trusted relationships with customers and partner with them for success.
  • Work on end-to-end services, spanning customer sites and platforms.
  • Collaborate with a talented team of professionals.
  • Embrace an entrepreneurial mindset and take ownership of responsibilities.
  • Seek innovative solutions.
  • Focus on quality, robustness, and security.
  • Implement cutting-edge tools to enhance operations, improve reliability, and gather feedback.
  • Identify and mitigate common operational issues to deliver seamless customer experiences.

Requirements

  • 10+ years of experience in operational management, including incident management and escalations.
  • Experience with design and implementation of application monitoring to ensure reliability and performance meets or exceeds business goals.
  • Experience implementing strategies to cap operations load and to handle overflow using appropriate tooling and metrics.
  • Experience defining service level indicators and objectives in collaboration with stakeholders, business, development, DevSecOps and Operations teams.
  • Solution and design experience in an enterprise environment: Windows server, Linux server (RHEL is preferred), UNIX (AIX, Solaris).
  • Experience with Windows server, storage, and Hyperscaler Cloud (AWS, Azure, Google Cloud Platform).
  • Experience with public cloud platforms such as AWS, OpenShift, Azure or GCP.
  • Experience working with Data format and Scripting languages JSON, YAML, Bash and/or PowerShell.

Skills

  • Operational management
  • Incident management
  • Escalations
  • Application monitoring
  • Reliability
  • Performance
  • Operations load capping
  • Overflow handling
  • Tooling
  • Metrics
  • Service level indicators (SLIs)
  • Service level objectives (SLOs)
  • Stakeholder collaboration
  • Business collaboration
  • Development collaboration
  • DevSecOps collaboration
  • Operations collaboration
  • Solution design
  • Enterprise environment design
  • Windows server
  • Linux server
  • RHEL
  • UNIX
  • AIX
  • Solaris
  • Storage
  • Hyperscaler Cloud
  • AWS
  • Azure
  • Google Cloud Platform
  • Public cloud platforms
  • OpenShift
  • Data format
  • JSON
  • YAML
  • Bash
  • PowerShell
  • Ansible
  • Terraform
  • Python
  • Distributed technologies
  • Dynamic resource management frameworks
  • Kubernetes
  • Open-source tooling
  • Prometheus
  • Grafana
  • Loki

Location

  • Global

Work Type

  • Hybrid

Experience Level

  • 10+ years of experience in operational management

Education Level

  • BS degree in Computer Science, Engineering, or other highly technical, scientific discipline

Benefits

  • Be Well programs designed to support financial, mental, physical, and social health.
  • Access to cutting-edge learning opportunities, including certifications with Microsoft, Google, and Amazon.
  • Coaching and hands-on experiences.

About the Company

  • Kyndryl runs and reimagines mission-critical technology systems that drive advantage for leading businesses.
  • We are at the heart of progress, enabling smarter decisions, faster innovation, and a lasting competitive edge through proven expertise and AI-powered insight.
  • Kyndryl has a global footprint.
  • Our culture is built on kinship, focusing on ensuring all Kyndryls feel included and welcoming people of all cultures, backgrounds, and experiences.
  • We believe in growth and are excited to see what candidates can bring.
  • A sense of belonging—being a valued, respected, trusted member of the team—is fundamental to our culture.
  • Kyndryl is dedicated to welcoming everyone, enabling individuals to thrive and contribute to a culture of empathy and shared success.
  • We achieve progress the world depends on, with purpose.
  • We are committed to sustainable progress for our customers and supporting the communities where we work and live.

Equal Opportunity

  • We welcome people of all cultures, backgrounds, and experiences.
  • Even if you don’t meet every requirement, we encourage you to apply.
  • Kyndryl gives you the ability to thrive and contribute to our culture of empathy and shared success.