Sr. Manager, Site Reliability at Omnicell | United States | Rezi

Sr. Manager, Site Reliability at Omnicell

Sr. Manager, Site Reliability

Omnicell · United States

1 weeks ago

Sr. Manager, Site Reliability

Omnicell · United States

14 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Omnicell is establishing a Global Cloud Operations organization to support its shift to a cloud-native, SaaS-delivered platform. This Site Reliability Engineering role is the first senior hire, responsible for designing the practice, setting standards, and leading initial operations. The role involves defining reliability practices, including SLOs, incident management, on-call rotations, observability platforms, and prioritizing reliability investments. It also anchors the company's investment in AI-driven operations, focusing on integrating AIOps and ML-assisted observability into monitoring and response processes within a regulated environment (HIPAA, SOC 2, FedRAMP).

Responsibilities

  • Define and publish SLOs and SLIs for top Tier-1 customer-facing services.
  • Establish error budget policy and enforcement mechanisms.
  • Design the incident command structure, including severity rubric, declaration criteria, and communication protocols.
  • Train incident commanders across Engineering and Support.
  • Select and implement the primary observability platform.
  • Define instrumentation standards for new services.
  • Partner to migrate incident response RACI into a durable SRE-owned model.
  • Establish the on-call rotation model, including compensation and handoff protocols.
  • Develop and track operational KPIs (MTTR, SLO attainment, etc.).
  • Present reliability metrics and roadmaps to senior leadership.
  • Instrument Tier-1 services, create dashboards, alerts, and runbooks.
  • Take the pager and command Sev-1 and Sev-2 incidents.
  • Lead blameless postmortems and drive follow-up work.
  • Contribute code and infrastructure-as-code (Terraform preferred).
  • Oversee the design and evolution of CI/CD pipelines.
  • Administer and scale Kubernetes platform, ensuring secure and compliant configurations.
  • Run chaos and failover exercises to validate resilience.
  • Architect Omnicell's AIOps direction, evaluating and introducing ML-based tools.
  • Integrate AI-assisted tooling into the observability and incident response stack.
  • Ensure AI-assisted operations meet regulatory and explainability requirements.
  • Coach one Engineer III SRE, providing guidance on incidents and design proposals.
  • Design future SRE hires, write requisitions, and run interview loops.
  • Represent SRE in architecture reviews and readiness reviews.
  • Partner with Enterprise Security and Compliance to ensure platform services meet regulatory requirements.

Requirements

  • Proven experience leading SRE, DevOps, or platform engineering teams in a cloud-native production environment.
  • Demonstrated experience building an SRE practice from zero or near-zero.
  • Deep hands-on expertise with at least one major public cloud (AWS, Azure, or GCP), including networking, IAM, and managed services.
  • Strong background in CI/CD pipeline design and management.
  • Experience implementing Infrastructure as Code using Terraform (preferred), Chef, Puppet, or similar tools.
  • Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.
  • Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations.
  • Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd).
  • Hands-on experience designing modern observability platforms.
  • Familiarity with integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage pipelines.
  • Real incident command experience for customer-impacting Sev-1 events, with blameless postmortem practice.
  • Ability to coach and mentor junior and mid-level engineers.
  • Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.
  • Excellent communication and stakeholder management skills.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent Experience.
  • 8+ years of experience in software or platform engineering, with at least 4 in an SRE, DevOps, or platform reliability role.
  • Proven experience advising and influencing senior technical or operations leaders.
  • At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.

Skills

  • SRE
  • DevOps
  • Platform Engineering
  • Cloud-Native
  • Production Environment Management
  • SLOs
  • Incident Command
  • Error Budgets
  • AWS
  • Azure
  • GCP
  • Networking
  • IAM
  • Managed Services
  • CI/CD
  • CodeFresh
  • GitHub Actions
  • Jenkins
  • TeamCity
  • Infrastructure as Code
  • Terraform
  • Chef
  • Puppet
  • Python
  • Object-Oriented Programming
  • Kubernetes Administration
  • Docker
  • Helm
  • Service Mesh
  • Istio
  • Linkerd
  • Observability Platforms
  • DataDog
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Elasticsearch
  • Kibana
  • AI/ML
  • Anomaly Detection
  • Alerting
  • LLM
  • Incident Response
  • Blameless Postmortems
  • Coaching
  • Mentoring
  • HIPAA
  • SOC 2
  • Communication
  • Stakeholder Management
  • Technical Leadership

Location

  • Remote
  • Hybrid

Work Type

  • Hybrid
  • Remote
  • On-call

Experience Level

  • Senior
  • 8+ years of experience
  • 4+ years in SRE/DevOps/Reliability
  • 2+ years technical leadership/mentorship

Education Level

  • Bachelor's degree in Computer Science, Engineering, or related technical field OR equivalent Experience
  • Masters Degree

About the Company

  • Omnicell is shifting from on-premise, hardware-centric products to a cloud-native, SaaS-delivered platform that hospitals depend on 24/7.
  • The products Omnicell builds dispense medications in hospitals, directly impacting patient care.
  • Operates in a regulated environment, adhering to HIPAA, SOC 2, and potentially FedRAMP.