About the Role
Kestra Holdings is seeking a Director of Observability to establish a new observability and reliability practice from the ground up. This is a hands-on leadership role responsible for selecting and configuring tools, defining instrumentation standards, building dashboards and alerting, and managing incidents during the initial phase. The Director will report to the Head of IT Infrastructure & Cybersecurity and will lead a growing team.
Responsibilities
- Define and execute the observability strategy for Kestra Holdings.
- Lead the initial build-out of the observability platform across metrics, logs, traces, profiles, and alerting.
- Work with and enforce existing instrumentation standards across infrastructure and application teams.
- Build the first generation of dashboards, SLO scorecards, and a single pane of glass for Tier-1 service health.
- Operate the firm's end-to-end incident management lifecycle.
- Stand up on-call schedules, escalation policies, and runbook-driven triage.
- Serve as primary incident commander for major incidents during the initial build phase.
- Integrate the incident lifecycle with Jira / Jira Service Management (JSM).
- Facilitate post-incident reviews (PIRs) and track remediation items.
- Drive adoption of SRE principles across the firm.
- Establish release of reliability gates and embed reliability into the service lifecycle.
- Partner with Cloud & Platform Engineering, Cybersecurity, and application teams to ensure services are instrumented and measurable.
- Ensure observability and incident management practices align with NIST CSF 2.0 maturity targets.
- Partner with the Cybersecurity team to integrate observability data with Jira/JSM, CMDB, and SIEM.
- Support regulatory and audit requirements for a SEC-regulated financial services firm.
- Directly lead and mentor a team of two: a Senior Observability Architect and an Observability/Reliability Engineer.
- Operate as a player-coach, splitting time between strategic leadership, hands-on engineering, and mentorship.
- Build a multi-year workforce plan and talent pipeline to scale the team.
- Foster a culture of blameless learning, operational excellence, and engineering-led reliability.
- Serve as the primary technical liaison for observability and incident management vendors.
- Represent Observability & Reliability in the Architecture Review Board, IT Change Management Board, and incident command forums.
- Provide regular reporting to SVP and executive leadership on reliability KPIs and incident trends.
Requirements
- 10+ years in observability, SRE, platform engineering, or infrastructure operations roles.
- 3+ years in a people leadership capacity (Director or Sr. Manager level).
- Demonstrated experience building an observability or SRE practice from scratch.
- Willing and able to write code/IaC, configure platforms, build dashboards, and run incidents personally.
- Deep expertise across the observability stack: metrics, log aggregation, distributed tracing, and profiling.
- Proven experience defining SLIs/SLOs, error budgets, and toil reduction programs.
- Hands-on experience with incident management platforms and Jira / Jira Service Management integration.
- Experience leading distributed teams across US and India time zones.
- Experience operating in a regulated industry (financial services, healthcare, or similar).
- Familiarity with compliance frameworks (NIST CSF, SOC 2, SEC, FINRA).
- Excellent communication skills, able to present reliability posture and risk to executive audiences.
- Experience with IaC (Terraform, Bicep, ARM), CI/CD pipelines, and embedding observability-as-code.
- Internal applicants must be in good standing with a minimum of 1 year of service with Kestra and 1 year in their current role unless approved by EVP.
Skills
- Observability
- SRE
- Platform Engineering
- Infrastructure Operations
- People Leadership
- Tool Selection
- Instrumentation Rollout
- SLO Definition
- Incident Command
- IaC (Terraform, Bicep, ARM)
- CI/CD Pipelines
- DevOps
- Metrics (Prometheus, Datadog, Azure Monitor)
- Log Aggregation (Elastic/OpenSearch, Log Analytics, Splunk)
- Distributed Tracing (OpenTelemetry, Jaeger, Datadog APM)
- Profiling
- Incident Management Platforms (PagerDuty, xMatters)
- Jira / Jira Service Management
- NIST CSF
- SOC 2
- SEC
- FINRA
- Azure Monitor/Log Analytics
- Datadog
- Grafana
- OpenTelemetry
- Elastic/Splunk
- PagerDuty/xMatters
- Jira / Jira Service Management (JSM)
Location
- Nationwide
Work Type
- Full-time
Experience Level
- Director
- Senior Manager
- 10+ years experience
- 3+ years leadership experience
Benefits
- Competitive pay and benefits
- 401(k)
- Health insurance
- Supportive, collaborative environment
- Opportunities for training, development, and long-term growth
- Tuition reimbursement
About the Company
- Kestra Holdings offers industry-leading wealth management platforms for independent wealth management professionals nationwide.
- Kestra is dedicated to empowering independent financial professionals—including traditional and hybrid RIAs—to grow their businesses and deliver exceptional client service.
- We combine advanced business management technology with personalized consulting to provide unmatched scale, efficiency, and support.
- Our advisor-focused culture is built on innovation and advocacy, enabling advisors to offer comprehensive securities and investment advisory solutions to their clients.
- Our Mission is Powering Financial Independence, enabling the growth and success of investing clients and the advisors who serve them.
- We live our values: Serve, Make it Happen, and One team.
Equal Opportunity
- It is the policy of Kestra Financial to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, sex, sexual orientation, gender, identity or expression, age, disability, marital status, citizenship, national origin, genetic information, or any other characteristic protected by law. Kestra Financial prohibits any such discrimination or harassment.
