Site Reliability Engineer, Team Lead at Omnicell | Austin, Texas, US | Rezi

Site Reliability Engineer, Team Lead at Omnicell

Site Reliability Engineer, Team Lead

Omnicell · Austin, Texas, US

3 weeks ago

Site Reliability Engineer, Team Lead

Omnicell · Austin, Texas, US

21 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Omnicell is building a Global Cloud Operations organization from the ground up as our business shifts from on-premise, hardware-centric products to a cloud-native, SaaS-delivered platform that hospitals depend on 24/7. The Site Reliability Engineering function is the reliability engine of that organization, and this role is the first senior SRE hire — the person who will design the practice, set the standards, and then run the plays themselves until the team is large enough to delegate. This is not a role where reliability practices already exist and you tune them. It is a role where you define what good looks like for Omnicell: which services have SLOs and at what targets, how incidents are declared and commanded, what the on-call rotation feels like, which observability platform we standardize on, and how reliability investment is prioritized against feature velocity. You will make those calls in partnership with the VP of Global Cloud Operations and an Engineer III SRE you will coach and grow. The environment is hybrid. Some of our products are still hardware in hospitals communicating with cloud services; others are fully SaaS. Some customers access us over private circuits, others over the public internet. We operate in a regulated environment — HIPAA, SOC 2, and in some engagements FedRAMP — which means reliability, security, and auditability are not separable concerns. The person we hire will be comfortable with that complexity and will help the organization design for it rather than around it. This role also anchors Omnicell's forward investment in AI-driven operations. Over the course of the first year, the organization intends to incorporate AIOps and ML-assisted observability — anomaly detection, intelligent alert correlation, LLM-assisted runbook generation — into how we monitor and respond to our platform. You will be the technical owner of how that gets introduced, prioritized against foundational reliability work, and validated in a regulated environment.

Responsibilities

  • Define and publish SLIs, SLOs, and error budgets for the top 5–10 Tier‑1 customer‑facing services in partnership with Product and Engineering.
  • Design Omnicell’s incident command structure, including severity definitions, declaration criteria, war‑room protocols, stakeholder communications, and post‑incident review standards.
  • Establish and operationalize a sustainable on‑call model, including fair rotations, paging discipline, escalation paths, and coordination with managed service partners (IBM, HCL).
  • Partner with the VP to migrate the interim incident response RACI — currently held by matrixed individuals across IT, Engineering, Support, and Enterprise Security — into a durable SRE-owned model.
  • Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement.
  • Define the instrumentation standards all new services must meet.
  • Develop and track operational KPIs (e.g., MTTR, SLO attainment, change‑failure rate, incident recurrence, cost per workload) and present reliability insights and roadmaps in executive Cloud Ops reviews.
  • Instrument Tier‑1 services directly—building dashboards, alerts, and runbooks yourself.
  • Participate in on‑call rotations and command Sev‑1 and Sev‑2 incidents, leading blameless postmortems and driving corrective actions to completion.
  • Contribute production code and infrastructure‑as‑code (Terraform preferred) to the platform.
  • Oversee the design and evolution of the CI/CD pipelines.
  • Administer and scale our Kubernetes platform, including secure and compliant cluster configurations.
  • Plan and execute chaos and failover exercises to validate real‑world resilience.
  • Architect Omnicell’s AIOps strategy, evaluating ML‑based anomaly detection, alert correlation, automated root‑cause analysis, and LLM‑assisted runbooks.
  • Make disciplined build‑versus‑buy decisions and integrate AI tooling only where it delivers measurable reliability gains.
  • Ensure AI‑assisted operations meet auditability, explainability, and compliance requirements (HIPAA, SOC 2).
  • Serve as formal coach to an Engineer III SRE, pairing on incidents, reviewing designs proposals, and supporting growth toward senior levels.
  • Design the next 2–4 SRE hires, including role definitions, interview loops, and hiring decisions.
  • Represent SRE in architecture reviews, launch readiness assessments, and cross‑functional reliability discussions.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent experience
  • 7+ years of experience in software or platform engineering, with at least 4 of those in an SRE, DevOps, or platform reliability role.
  • At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.
  • Proven experience leading SRE, DevOps, or platform engineering teams in a cloud-native production environment — with demonstrated experience building a practice from zero or near-zero: you have set SLOs, defined incident command, and introduced error budget thinking to an organization that did not have it.
  • Deep hands-on expertise with at least one major public cloud (AWS, Azure, or GCP), including networking, IAM, and managed services.
  • Strong background in CI/CD pipeline design and management (familiarity with CodeFresh, GitHub Actions, Jenkins, TeamCity, or equivalent).
  • Experience implementing Infrastructure as Code using Terraform (preferred), Chef, Puppet, or similar tools.
  • Proficiency in Python or another object-oriented programming language for automation, tooling, and production services.
  • Experience administering and scaling Kubernetes clusters, including secure and compliant platform configurations.
  • Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like.
  • Familiarity with integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage pipelines — or strong conviction about where AIOps should and should not be applied in a regulated environment.
  • Real incident command experience for customer-impacting Sev-1 events, with blameless postmortem practice and documented follow-up discipline.
  • Ability to coach and mentor, with direct evidence of growing junior and mid-level engineers.
  • Comfort operating in a regulated environment where reliability and compliance (HIPAA, SOC 2) are inseparable.
  • Collaborate: Partner deeply with Product, Platform Engineering, Support, Security, and managed service providers to align reliability with business priorities.
  • Inspire: Lead by example during high‑stakes incidents and influence teams toward a culture of ownership, learning, and resilience.
  • Develop: Invest in the growth of your SRE peers through coaching, pairing, and thoughtful technical leadership.
  • Execute: Set clear priorities, make informed trade‑offs, and deliver durable reliability improvements.
  • Impact: Shape how Omnicell operates for years to come by defining the standards, tools, and practices of our SRE function.
  • Modeling a growth mindset and continuous learning.
  • Acting as a talent activator through formal coaching and mentorship.
  • Being an impact maker who connects reliability investment to business and patient outcomes.
  • Serving as a change champion as Omnicell transitions to cloud‑first operations.

Skills

  • SRE
  • DevOps
  • Platform Reliability
  • Technical Leadership
  • Mentorship
  • Cloud-native production environment
  • AWS
  • Azure
  • GCP
  • Networking
  • IAM
  • Managed Services
  • CI/CD pipeline design
  • CodeFresh
  • GitHub Actions
  • Jenkins
  • TeamCity
  • Infrastructure as Code
  • Terraform
  • Chef
  • Puppet
  • Python
  • Object-oriented programming
  • Kubernetes
  • Docker
  • Helm
  • Service Mesh
  • Istio
  • Linkerd
  • DataDog
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Elasticsearch
  • Kibana
  • AI/ML-based anomaly detection
  • Alerting
  • LLM-assisted triage
  • Incident Command
  • Blameless Postmortem
  • HIPAA
  • SOC 2

Location

  • Austin, TX
  • Cranberry Woods, PA
  • Remote (USA)

Work Type

  • Hybrid
  • Remote

Experience Level

  • Senior
  • Team Lead
  • Player-Coach

Education Level

  • Bachelor's degree in Computer Science, Engineering, or a related technical field OR equivalent experience

About the Company

  • Since 1992, Omnicell has been committed to transforming pharmacy care through outcomes-centric innovation designed to optimize clinical and business outcomes across all settings of care.
  • We strive to be the healthcare provider’s most trusted partner by our guiding promise of “Outcomes. Defined and Delivered.”
  • Our comprehensive portfolio of robotics, smart devices, intelligent software, and expert services is helping healthcare facilities worldwide to improve business and clinical outcomes as they move closer to the industry vision of the Autonomous Pharmacy.
  • Our guiding principles inform everything we do: As Passionate Transformers, we find a better way to innovate relentlessly.
  • Being Mission Driven, we consistently deliver on our promises.
  • Our Entrepreneurial spirit makes the most of EVERY opportunity for innovation.
  • Understanding that Relationships Matter creates synergies that yield the greatest benefits for all.
  • Intellectually Curious, eager to think deeper to learn and improve.
  • In Doing the Right Thing, we lead by example in ALL we do.
  • We are deeply committed to Environmental, Social, and Governance (ESG) initiatives.
  • Our ESG efforts focus on creating an inclusive culture and a healthier world.
  • This includes our Employee Impact Groups, which foster inclusion and belonging, as well as our learning and well-being programs that support personal and professional growth.
  • We also prioritize sustainability in our operations, aiming to reduce our environmental footprint and promote responsible business practices.
  • Join us in transforming the pharmacy care delivery model, making patient care safer and smarter for all.

Equal Opportunity

  • Omnicell is dedicated to fostering an inclusive workplace.
  • We welcome applications from all individuals, valuing a wide range of perspectives and backgrounds.
  • As an equal opportunity employer, we do not discriminate based on race, gender, religion, sexual orientation, gender identity, national origin, veteran status, or disability.
  • We are committed to making our recruitment process accessible to everyone.
  • We offer support and reasonable adjustments for individuals with disabilities during our hiring process.
  • If you need assistance, please contact us at Recruiting@omnicell.com [Recruiting@omnicell.com].
  • At Omnicell, respect for privacy and confidentiality is paramount.
  • We adhere to strict policies to prevent discrimination or retaliation against those who engage in open conversations about compensation.
  • However, employees privy to compensation information as part of their job role are expected to maintain confidentiality, except in specific circumstances outlined by law, such as during formal complaints, investigations, or as required by legal obligations.
  • Please note that Omnicell reserves the right to modify job roles and responsibilities as needed to meet our organization's evolving needs and drive our mission forward.