Principal Site Reliability Engineer, Infrastructure Observability at T. Rowe Price | GB | Rezi

Principal Site Reliability Engineer, Infrastructure Observability at T. Rowe Price

Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price · GB

1 weeks ago

Principal Site Reliability Engineer, Infrastructure Observability

T. Rowe Price · GB

8 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

In this role, you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability, and recoverability of T. Rowe Price’s innovative cloud & on-prem solutions by leveraging automation and best-of-breed tools. The successful candidate will have a strong operations & engineering background, be hands-on when needed, and possess expertise in cloud environments, infrastructure operations, DevOps practices, CI/CD toolchain, code build and deployment, incident response, and 24x7 monitoring and support. You will also have extensive experience operating within a SRE function in a complex, distributed environment and a demonstrated ability to work horizontally and vertically within an organization with diverse partners and sponsor groups.

Responsibilities

  • Design technology solutions to prevent or minimize service disruptions.
  • Prevent technology service disruptions through technology solution recommendations and automations.
  • Foster a culture of deep learning through blameless post-mortems to improve the shared goal of reliability across services.
  • Transform operations teams by facilitating internal change to adopt SRE standard methodologies and driving strategic growth in this area.
  • Analyze incidents impacting technology availability for high-level trends across the broad portfolio.
  • Drive initiatives to reduce or prevent technology failures in a complex, distributed technology environment.
  • Consolidate information from disconnected systems into cohesive views of the technology portfolio for identifying trends, redundancies, and risk.
  • Lead initiatives of varying complexity that span multi-functional areas.
  • Contribute to the definition of target state architecture and design of the technology environment.

Requirements

  • Bachelor's degree or equivalent combination of education and relevant experience.
  • 10+ years of experience designing and operating cloud infrastructure with senior-level impact.
  • 5+ years building and supporting solutions in Amazon AWS.
  • 5+ years of experience building and running a DevOps and/or SRE function.
  • Experience with implementation and operation of the chaos model at scale.
  • Strategic and program-level implementation experience.
  • Demonstrable experience implementing new technology, tools, and platforms.
  • System administration and scripting experience.
  • Demonstrable experience leveraging automation to proactively prevent or quickly remediate incidents.
  • Fluent in multiple programming languages (e.g., Python, Java, GO, Node.js, .Net Core, etc).
  • Proficiency with database development (SQL Server, PostgreSQL, MySQL, etc).
  • Proficiency with defining, right-sizing, tracking, and reporting on Service Level Objectives (SLOs), Service Level Indicators (SLIs), system availability, and the progress and outcomes related to reliability.
  • Experience with implementing and managing Error Budgets.
  • Proficiency with understanding and explaining incident situations and their recovery plans to prevent recurrence.
  • Knowledge/experience driving dashboard standardization across the ecosystem for observability, APM and infrastructure monitoring, and application-specific logging.
  • Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, and cloud native tools.
  • Knowledge/experience with cloud management tools such as Ansible, Terraform, Vault, and Vagrant.
  • Ability to work independently with guidance only in the most complex situations.
  • Ability to make sound decisions with limited facts or resources.
  • Ability to balance strategic and pragmatic concerns when solving problems.
  • Ability to adjust communication style and materials to suit a given audience.
  • Ability to clearly articulate operational principles, practices, and policies.
  • Ability to stay abreast of industry trends and technologies.
  • Accountability for work of self and others; setting standards for others.
  • Ability to maintain a broad internal professional network and know when to engage/activate it.
  • Ability to develop or mentor diverse talent on the team.
  • Ability to be on-call and/or work during off-hours.
  • Cloud or SRE-related certifications (Preferred).
  • Working knowledge of Azure (Preferred).

Skills

  • Observability
  • Sustainability
  • Scalability
  • Measurability
  • Recoverability
  • Automation
  • Cloud environments (public, private)
  • Infrastructure operations
  • DevOps practices
  • CI/CD toolchain
  • Systems
  • Code build and deployment
  • Incident response
  • 24x7 monitoring and support
  • Python
  • Java
  • GO
  • Node.js
  • .Net Core
  • SQL Server
  • PostgreSQL
  • MySQL
  • Service Level Objectives (SLOs)
  • Service Level Indicators (SLIs)
  • Error Budgets
  • New Relic
  • SolarWinds DPA
  • Elastic Stack
  • Prometheus
  • Grafana
  • Splunk
  • Ansible
  • Terraform
  • Vault
  • Vagrant

Location

  • Hybrid

Work Type

  • Hybrid
  • Full-time

Experience Level

  • Principal
  • Senior-level
  • 10+ years
  • 5+ years

Education Level

  • Bachelor's degree

About the Company

  • T. Rowe Price is an investment management firm.