Site Reliability Engineer at Datacom | AU | Rezi

Site Reliability Engineer at Datacom

Site Reliability Engineer

Datacom · AU

1 months ago

Site Reliability Engineer

Datacom · AU

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a motivated and technically capable Site Reliability Engineer to help maintain, and continuously improve the reliability, scalability, security, and performance of our enterprise platforms. You will work across cloud technologies, infrastructure automation, observability, security, and operational processes to ensure services remain resilient and available. This role requires a strong operational mindset, a passion for automation, and the ability to troubleshoot complex technical issues across hybrid and cloud environments.

Responsibilities

  • Implement, and maintain highly available and resilient infrastructure.
  • Drive continuous service improvements through performance analysis and operational metrics.
  • Support disaster recovery planning including Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO).
  • Develop and maintain monitoring, logging, and alerting solutions.
  • Proactively identify reliability risks before they impact services.
  • Analyse system performance and capacity trends.
  • Create operational dashboards that provide actionable insights.
  • Participate in and support on-call rotations.
  • Lead troubleshooting and resolution of production incidents and service outages.
  • Develop and maintain incident response runbooks and playbooks.
  • Conduct blameless post-incident reviews and drive corrective actions.
  • Implement and maintain Infrastructure as Code (IaC) solutions.
  • Automate operational tasks, deployments, and service recovery processes.
  • Embed security controls throughout the technology lifecycle using DevSecOps principles.
  • Contribute to CI/CD pipeline development and optimisation.
  • Support cloud environments across Azure, AWS and Google Cloud Platform.
  • Assist with cloud architecture, governance, landing zones, and platform standardisation.
  • Support containerised workloads using Docker and Kubernetes.
  • Collaborate with engineering and security teams to improve platform reliability and performance.

Requirements

  • Must have full working rights (Citizen, Permanent Resident or Visa Holder).
  • No sponsorships provided.
  • Able to get a National Police Check clearance.
  • 5+ years in Infrastructure, Cloud Engineering, Platform Engineering, Systems Engineering, or Site Reliability Engineering roles.
  • Experience supporting enterprise cloud platforms including: Microsoft Azure, Amazon Web Services (AWS), Google Cloud Platform (GCP).
  • Experience administering: Active Directory, Microsoft Entra ID, Microsoft Intune, Google Workspace.
  • Strong understanding of virtualisation technologies.
  • Experience with cloud landing zones and platform governance.
  • Familiarity with containerisation technologies such as Docker and Kubernetes.
  • Experience implementing monitoring, logging and alerting solutions.
  • Understanding of SLIs, SLOs and Error Budgets.
  • Experience supporting high-availability environments.
  • Knowledge of incident management and problem management practices.
  • Understanding of business continuity, disaster recovery, RTO and RPO requirements.
  • Familiarity with Infrastructure as Code (IaC).
  • Familiarity with DevSecOps practices.
  • Familiarity with Zero Trust security principles.
  • Familiarity with Security and governance frameworks.
  • Working knowledge of NIST Cybersecurity Framework.
  • Working knowledge of ITIL.
  • Working knowledge of COBIT.
  • Working knowledge of ISO 27001.
  • Working knowledge of Zero Trust Architecture.

Skills

  • Microsoft Azure
  • Amazon Web Services (AWS)
  • Google Cloud Platform (GCP)
  • Active Directory
  • Microsoft Entra ID
  • Microsoft Intune
  • Google Workspace
  • Virtualisation technologies
  • Cloud landing zones
  • Platform governance
  • Docker
  • Kubernetes
  • Monitoring
  • Logging
  • Alerting solutions
  • SLIs
  • SLOs
  • Error Budgets
  • Incident management
  • Problem management
  • Business continuity
  • Disaster recovery
  • RTO
  • RPO
  • Infrastructure as Code (IaC)
  • DevSecOps
  • Zero Trust security principles
  • Security and governance frameworks
  • NIST Cybersecurity Framework
  • ITIL
  • COBIT
  • ISO 27001
  • Zero Trust Architecture

Location

  • Remote

Work Type

  • Remote working
  • Flexi-hours

Experience Level

  • 5+ years

Benefits

  • Social events
  • Chill-out spaces
  • Remote working
  • Flexi-hours
  • Professional development courses
  • Retail discounts

About the Company

  • Datacom connects people and technology in order to solve challenges, create opportunities and discover new possibilities for the communities we live in.
  • Datacom is one of Australia and New Zealand’s largest suppliers of Information Technology professional services.
  • We have managed to maintain a dynamic, agile, small business feel that is often diluted in larger organisations of our size.
  • It's our people that give Datacom its unique culture and energy that you can feel from the moment you meet with us.
  • We operate at the forefront of technology to help Australia and New Zealand’s largest enterprise organisations explore possibilities and solve their greatest challenges, so you will never run out of interesting new challenges and opportunities.
  • We want Datacom to be an inclusive and welcoming workplace for everyone and take pride in the steps we have taken and continue to take to make our environment fun and friendly, and our people feel supported.

Equal Opportunity

  • We want Datacom to be an inclusive and welcoming workplace for everyone and take pride in the steps we have taken and continue to take to make our environment fun and friendly, and our people feel supported.