Staff Operational Support Engineer at Dolby Laboratories, Inc. | Sydney, New South Wales | Rezi

Staff Operational Support Engineer at Dolby Laboratories, Inc.

Staff Operational Support Engineer

Dolby Laboratories, Inc. · Sydney, New South Wales

2 months ago

Staff Operational Support Engineer

Dolby Laboratories, Inc. · Sydney, New South Wales

2 months ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Dolby OptiView is establishing an Operational Support (L2) team to ensure the stability and excellence of its 24/7 live video streaming, ads, player, and real-time delivery platforms. As an Operational Support Engineer (L2), you will own customer-impacting production incidents from triage to resolution, operating directly on production systems and leading incident response. This role bridges Support, Engineering, DevOps, and customers, focusing on incident ownership, production operations, automation, and scalability.

Responsibilities

  • Own escalated customer issues from Level 1 Support and drive them to resolution.
  • Troubleshoot and resolve complex, high-impact production incidents affecting live streams, VOD playback, ad insertion, DRM, and real-time WebRTC services.
  • Operate directly on production environments, including configuration changes and CDN adjustments, following established procedures and executing emergency changes when necessary.
  • Lead or contribute to live incident bridges involving customers, internal teams, and partners.
  • Provide clear and timely communication during incidents, including status updates and customer-facing explanations.
  • Understand, troubleshoot, and safely modify production environments using Infrastructure as Code (IaC).
  • Use IaC (Terraform, Helm, Kubernetes manifests, GitOps) as the primary mechanism for operational changes.
  • Collaborate with Engineering and DevOps to improve deployment reliability and operational safety.
  • Validate and execute infrastructure or configuration changes through codified workflows.
  • Leverage AI tools and automation for enhanced operational efficiency and incident response.
  • Contribute to and use AI-assisted incident triage, automated runbook execution, AI-based pattern detection, and intelligent alert correlation.
  • Use AI to generate or improve incident communications, accelerate troubleshooting, and identify recurring issues.
  • Drive the adoption of automation-first and AI-augmented operational practices.
  • Participate in pre-event readiness planning for critical customer events.
  • Validate system readiness through runbook checks, monitoring coverage validation, and risk identification.
  • Define and rehearse incident response strategies for high-risk scenarios.
  • Collaborate with customers and internal teams to ensure smooth event execution.
  • Participate in a 24/7 on-call rotation, responding to critical alerts within defined SLAs.
  • Ensure smooth handovers between shifts and regions.
  • Perform or contribute to root cause analysis (RCA) for production incidents.
  • Document findings, corrective actions, and preventive measures.
  • Identify recurring issues and work with Engineering and Product teams to eliminate them.
  • Contribute to and improve runbooks, operational playbooks, and knowledge bases.
  • Work closely with Engineering teams to escalate defects, validate fixes, and support production deployments.
  • Provide feedback on system observability, tooling gaps, and operational risks.
  • Act as the operational voice during post-incident reviews.

Requirements

  • 5+ years of relevant experience in operational, support, or similar customer-facing roles.
  • Proven ability to own complex problems end-to-end and operate with a high degree of autonomy.
  • Strong experience supporting production video streaming platforms, OTT services, and live systems.
  • Solid troubleshooting skills across distributed systems (APIs, microservices, cloud infrastructure).
  • Comfort performing controlled changes in production environments.
  • Working knowledge of incident management and on-call operations.
  • Proven ability to remain calm, structured, and decisive during high-pressure incidents.
  • Strong sense of ownership and accountability for customer outcomes.
  • Excellent written and verbal communication skills, including customer-facing communication during incidents.

Skills

  • HLS
  • DASH
  • CMAF
  • WebRTC
  • DRM
  • CDN architectures
  • Grafana
  • Kibana/ELK
  • Prometheus
  • Loki
  • Terraform
  • Helm
  • Kubernetes manifests
  • GitOps workflows
  • CI/CD and deployment pipelines
  • AI tools and automation

Work Type

  • 24/7

Experience Level

  • 5+ years

About the Company

  • Dolby OptiView is building a dedicated Operational Support (L2) team responsible for the stability, availability, and operational excellence of our 24/7 live video streaming, ads, player, and real‑time delivery platforms.