Sr. Manager, Site Reliability at Omnicell | Austin, USA | Rezi

Sr. Manager, Site Reliability at Omnicell

Sr. Manager, Site Reliability

Omnicell · Austin, USA

2 weeks ago

Sr. Manager, Site Reliability

Omnicell · Austin, USA

15 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Omnicell is establishing a Global Cloud Operations organization to support its shift to a cloud-native, SaaS platform. This Site Reliability Engineering role is the first senior hire, responsible for designing the practice, setting standards, and leading initial operations. The role involves defining reliability practices, including SLOs, incident management, observability, and prioritization, in partnership with leadership. The environment is hybrid and regulated (HIPAA, SOC 2, FedRAMP), requiring comfort with complexity. The role also anchors the company's investment in AI-driven operations, with the technical owner responsible for integrating AIOps and ML-assisted observability.

Responsibilities

  • Define and publish SLOs and SLIs for top Tier-1 customer-facing services.
  • Establish error budget policy and enforcement mechanisms.
  • Design the incident command structure, including severity rubric, declaration criteria, and postmortem templates.
  • Train incident commanders across Engineering and Support.
  • Select and stand up the primary observability platform and define instrumentation standards.
  • Partner to migrate the interim incident response RACI into an SRE-owned model.
  • Establish the on-call rotation model, including compensation, paging discipline, and handoff protocols.
  • Develop and track operational KPIs and present reliability metrics to senior leadership.
  • Instrument Tier-1 services, write dashboards, alerts, and runbooks.
  • Take the pager, command Sev-1 and Sev-2 incidents, and lead blameless postmortems.
  • Contribute code and infrastructure-as-code to the platform.
  • Oversee the design and evolution of CI/CD pipelines.
  • Administer and scale Kubernetes platforms, ensuring secure and compliant configurations.
  • Run chaos and failover exercises to validate resilience.
  • Architect Omnicell's AIOps direction, evaluating and introducing ML-based anomaly detection and other AI tools.
  • Make build-versus-buy decisions for AIOps tooling and integrate AI-assisted tools where valuable.
  • Ensure AI-assisted operations meet auditability and explainability requirements in a regulated environment.
  • Coach one Engineer III SRE, providing guidance on incidents and design proposals.
  • Design the next 2-4 SRE hires, writing requisitions and running interview loops.
  • Represent SRE in architecture reviews, product launch readiness reviews, and executive metric reviews.
  • Partner with Enterprise Security, Compliance, and Architecture to ensure regulatory and security requirements are met.

Requirements

  • Proven experience leading SRE, DevOps, or platform engineering teams in a cloud-native production environment.
  • Demonstrated experience building an SRE practice from zero or near-zero.
  • Experience setting SLOs, defining incident command, and introducing error budget thinking.
  • Deep hands-on expertise with at least one major public cloud (AWS, Azure, or GCP), including networking, IAM, and managed services.
  • Strong background in CI/CD pipeline design and management.
  • Experience implementing Infrastructure as Code using Terraform (preferred), Chef, Puppet, or similar tools.
  • Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.
  • Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations.
  • Working knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd).
  • Hands-on experience designing modern observability platforms.
  • Familiarity with integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage pipelines.
  • Real incident command experience for customer-impacting Sev-1 events, with blameless postmortem practice and documented follow-up discipline.
  • Ability to coach and mentor, with direct evidence of growing junior and mid-level engineers.
  • Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.
  • Excellent communication and stakeholder management skills.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent Experience.
  • 8+ years of experience in software or platform engineering, with at least 4 of those in an SRE, DevOps, or platform reliability role.
  • Proven experience advising and influencing senior technical or operations leaders using data-driven recommendations.
  • At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.

Skills

  • SRE
  • DevOps
  • Platform Engineering
  • Cloud-Native
  • Production Environment
  • SLOs
  • Incident Command
  • Error Budgets
  • AWS
  • Azure
  • GCP
  • Networking
  • IAM
  • Managed Services
  • CI/CD
  • CodeFresh
  • GitHub Actions
  • Jenkins
  • TeamCity
  • Infrastructure as Code
  • Terraform
  • Chef
  • Puppet
  • Python
  • Object-Oriented Programming
  • Kubernetes
  • Docker
  • Helm
  • Service Mesh
  • Istio
  • Linkerd
  • Observability Platforms
  • DataDog
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Elasticsearch
  • Kibana
  • AI/ML
  • Anomaly Detection
  • Predictive Alerting
  • Automated Root Cause Analysis
  • LLM-assisted Triage
  • AIOps
  • Incident Response
  • Postmortems
  • Coaching
  • Mentorship
  • HIPAA
  • SOC 2
  • Communication
  • Stakeholder Management
  • Computer Science
  • Engineering
  • Technical Leadership
  • Data-Driven Recommendations
  • Masters Degree
  • Healthcare
  • Clinical Workflows
  • Regulated Vertical
  • Managed Service Providers
  • IBM
  • HCL
  • Hybrid Hardware-Plus-Cloud Products
  • Large Language Model APIs
  • Agentic AI Frameworks
  • Stateful Distributed Services
  • Security Scanning
  • Intrusion Detection Systems
  • Messaging Systems
  • Kafka
  • RabbitMQ
  • Chaos Engineering
  • Chaos Monkey
  • LitmusChaos
  • Databricks
  • Team Foundation Server
  • Octopus Deploy
  • FinOps
  • Cloud Cost Optimization

Location

  • Remote
  • Hybrid

Work Type

  • Hybrid
  • Remote
  • On-call participation

Experience Level

  • Senior
  • 8+ years of experience in software or platform engineering
  • 4+ years in SRE, DevOps, or platform reliability role
  • 2+ years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities

Education Level

  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent Experience
  • Masters Degree

About the Company

  • Since 1992, Omnicell has been committed to transforming pharmacy care through outcomes-centric innovation designed to optimize clinical and business outcomes across all settings of care.
  • We strive to be the healthcare provider’s most trusted partner by our guiding promise of “Outcomes. Defined and Delivered.”
  • Our comprehensive portfolio of robotics, smart devices, intelligent software, and expert services is helping healthcare facilities worldwide to improve business and clinical outcomes as they move closer to the industry vision of the Autonomous Pharmacy.
  • Our guiding principles inform everything we do: As Passionate Transformers, we find a better way to innovate relentlessly; Being Mission Driven, we consistently deliver on our promises; Our Entrepreneurial spirit makes the most of EVERY opportunity for innovation; Understanding that Relationships Matter creates synergies that yield the greatest benefits for all; Intellectually Curious, eager to think deeper to learn and improve; In Doing the Right Thing, we lead by example in ALL we do.
  • We are deeply committed to Environmental, Social, and Governance (ESG) initiatives.
  • Our ESG efforts focus on creating an inclusive culture and a healthier world.
  • This includes our Employee Impact Groups, which foster inclusion and belonging, as well as our learning and well-being programs that support personal and professional growth.
  • We also prioritize sustainability in our operations, aiming to reduce our environmental footprint and promote responsible business practices.
  • Join us in transforming the pharmacy care delivery model, making patient care safer and smarter for all.

Equal Opportunity

  • Omnicell is dedicated to fostering an inclusive workplace.
  • We welcome applications from all individuals, valuing a wide range of perspectives and backgrounds.
  • As an equal opportunity employer, we do not discriminate based on race, gender, religion, sexual orientation, gender identity, national origin, veteran status, or disability.
  • We are committed to making our recruitment process accessible to everyone.
  • We offer support and reasonable adjustments for individuals with disabilities during our hiring process.
  • If you need assistance, please contact us at Recruiting@omnicell.com.
  • At Omnicell, respect for privacy and confidentiality is paramount.
  • We adhere to strict policies to prevent discrimination or retaliation against those who engage in open conversations about compensation.
  • However, employees privy to compensation information as part of their job role are expected to maintain confidentiality, except in specific circumstances outlined by law, such as during formal complaints, investigations, or as required by legal obligations.
  • Please note that Omnicell reserves the right to modify job roles and responsibilities as needed to meet our organization's evolving needs and drive our mission forward.