Observability Engineer at Universal Music Group | London | Rezi

Observability Engineer at Universal Music Group

Observability Engineer

Universal Music Group · London

Today

Observability Engineer

Universal Music Group · London

8 hours ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

We are seeking a talented and proactive Observability Engineer to join our dynamic global team. You will be passionate about data driven decisions, automation, and committed to continuous improvement. In this pivotal role, you will be instrumental in building, maintaining, and enhancing the comprehensive observability solutions that ensure the reliability, performance, and scalability of our critical IT systems and applications across the globe. You'll work at the intersection of technology and the vibrant world of music, providing deep insights that drive operational excellence and enable rapid response to any challenge.

Responsibilities

  • Design, implement, and continuously improve observability stack including monitoring, logging, tracing, and alerting systems across diverse cloud-native, on-premise, and hybrid environments.
  • Evaluate, select, and implement observability tools and platforms, and automate observability pipelines and alerting mechanisms.
  • Define, enforce, and advocate for observability standards and best practices across all engineering and operations teams.
  • Create and maintain monitoring solutions, dashboards, and automated alerts for real-time insights into system health, performance, and availability.
  • Utilize telemetry data for swift diagnosis and resolution of incidents, and conduct post-incident reviews.
  • Partner with other teams to drive positive change and influence best practice within the Observability space.
  • Embed observability throughout the entire technology lifecycle, empowering teams with insights.
  • Analyze telemetry data to identify and resolve performance bottlenecks, optimize resource allocation, and fine-tune configurations.
  • Contribute to compliance and security efforts through effective log management and integration with SIEM systems.
  • Undertake system analysis based on Observability notifications to troubleshoot complex issues.
  • Document and define processes and best practices.
  • Actively contribute to making the Observability team a positive and respectful place to work.

Requirements

  • 3+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a dedicated focus on observability.
  • Strong understanding and practical experience with monitoring, logging, and tracing systems.
  • Proficiency with industry-standard observability tools (e.g., Dynatrace, AWS Cloudwatch, Prometheus, Grafana, ELK Stack, Splunk, Logic Monitor).
  • Strong technical knowledge with major cloud platforms (AWS, Azure, or GCP).
  • Solid programming and scripting skills (e.g., Python, Go, Shell, JavaScript) for automation.
  • Understanding of distributed systems, microservices architectures, and cloud-native environments.
  • Experience with Docker/Kubernetes and DevOps principles.
  • Familiarity with CI/CD pipelines and automation tools (e.g., Ansible, Terraform).
  • Exceptional analytical and problem-solving abilities, with a proactive approach to tackling complex technical challenges.
  • Excellent communication, collaboration, and interpersonal skills, with the ability to clearly articulate technical concepts to diverse audiences.
  • Prior experience supporting critical business applications within a large-scale, global enterprise environment.
  • Awareness of security best practices and the ability to integrate security monitoring into observability processes.
  • Self-motivated with high degree of initiative and excellent follow-up skills.
  • Experience with Chaos Engineering, Canary/BlueGreen deployment strategies, capacity planning, data analysis, networking.
  • Experience in designing and automating Observability workloads with a foundation in software engineering, database administration and system administration.
  • Scripting and programming for Observability as well as troubleshooting across Python, Go, Java and associated languages.

Skills

  • Observability
  • Site Reliability Engineering (SRE)
  • DevOps
  • Monitoring
  • Logging
  • Tracing
  • Alerting
  • Dynatrace
  • AWS Cloudwatch
  • Prometheus
  • Grafana
  • ELK Stack
  • Splunk
  • Logic Monitor
  • AWS
  • Azure
  • GCP
  • Python
  • Go
  • Shell
  • JavaScript
  • Distributed systems
  • Microservices architectures
  • Cloud-native environments
  • Docker
  • Kubernetes
  • DevOps principles
  • CI/CD pipelines
  • Ansible
  • Terraform
  • Chaos Engineering
  • Canary/BlueGreen deployment strategies
  • Capacity planning
  • Data analysis
  • Networking
  • Software engineering
  • Database administration
  • System administration
  • Java

Location

  • Global

Work Type

  • Hybrid
  • Full-time

Experience Level

  • 3+ years

Education Level

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.

About the Company

  • Universal Music Group is the world’s leading music company.
  • We are committed to artistry, innovation and entrepreneurship.
  • We own and operate a broad array of businesses engaged in recorded music, music publishing, merchandising, and audiovisual content in more than 60 countries.
  • We identify and develop recording artists and songwriters, and we produce, distribute and promote the most critically acclaimed and commercially successful music to delight and entertain fans around the world.

Equal Opportunity

  • Everyone is welcome to apply for our roles, and we are determined to ensure that no applicant or employee receives less favourable treatment because of gender, race, disability, sexual orientation, religion, belief, age, marital status, background, pregnancy, or caring responsibilities.
  • We also recognise the importance of diversity of thought within our teams and are fully committed to embracing the talents of people with autism, dyslexia, ADHD, and other forms of neurocognitive variation.
  • We will always seek to make appropriate adjustments to recruitment, workplaces, and work processes to be fully inclusive to people with different needs and working styles.
  • If you need us to make any reasonable adjustments for you from application onwards, including alternatives to the online form or to disclose a neurocognitive condition, please email UniversalMusicCareers@umusic.com.