Engineering Manager, Site Reliability Engineering at Google | AU | Rezi

Engineering Manager, Site Reliability Engineering at Google

Engineering Manager, Site Reliability Engineering

Google · AU

1 weeks ago

Engineering Manager, Site Reliability Engineering

Google · AU

13 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Engineering Manager, Site Reliability Engineering role.

Rezi rewrites your resume against Google's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Engineering Manager, Site Reliability Engineering posting at Google — free, in seconds.

About the Role

Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services have reliability, uptime appropriate to customer's needs and a fast rate of improvement. SREs will keep an ever-watchful eye on our systems capacity and performance. This role involves optimizing existing systems, building infrastructure, and eliminating work through automation, while managing complex challenges of scale unique to Google Cloud. The SRE Manager will lead a team of engineers tasked with ensuring services are fast, reliable, and available, balancing new feature development with the long-term health and stability of the production environment.

Responsibilities

  • Lead and grow a team of SREs, providing technical guidance, career development, and performance management.
  • Partner with Software Development teams to design, build, and maintain scalable and reliable services.
  • Drive the evolution of monitoring, alerting, and incident response frameworks to ensure rapid detection and resolution of production issues.
  • Implement software solutions and automation to eliminate manual tasks and improve the efficiency of systems operations.
  • Manage Service Level Objectives (SLOs) and error budgets to balance innovation with system stability.
  • Participate in on-call rotations as a manager-escalation point and lead incident command during critical outages.
  • Advocate a culture of engineering excellence, technical curiosity, and continuous improvement within the Sydney office.

Requirements

  • Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
  • 8 years of experience with software development in one or more programming languages.
  • 3 years of experience managing people or teams.
  • 3 years of experience leading projects.
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems.
  • Master’s degree or PhD in Computer Science or a technical field.
  • Experience managing high-performing teams responsible for large-scale, 24/7 production systems.
  • Deep understanding of SRE principles, including error budgets, toil reduction, and capacity planning.
  • Proven track record of driving cross-functional projects and influencing architectural decisions.
  • Excellent communication and stakeholder management skills, with the ability to translate complex technical concepts for non-technical audiences.

Skills

  • Software development
  • Programming languages
  • People management
  • Team leadership
  • Project leadership
  • Distributed systems design
  • Distributed systems analysis
  • Distributed systems troubleshooting
  • SRE principles
  • Error budgets
  • Toil reduction
  • Capacity planning
  • Cross-functional project driving
  • Architectural decision influencing
  • Communication
  • Stakeholder management
  • Technical concept translation

Location

  • Sydney

Work Type

  • Full-time

Experience Level

  • 8 years of experience with software development
  • 3 years of experience managing people or teams
  • 3 years of experience leading projects
  • 3 years of experience designing, analyzing, and troubleshooting distributed systems

Education Level

  • Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
  • Master’s degree or PhD in Computer Science or a technical field.

About the Company

  • Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another.
  • Our Site Reliability Engineering (SRE) team combines software engineering and systems administration to build and operate some of the world's most complex distributed systems.
  • Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running.
  • From developing and maintaining our data centers to building the next generation of Google platforms, we make Google's product portfolio possible.
  • We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them.
  • We keep our networks up and running, ensuring our users have the best and fastest experience possible.