About the Role
We are seeking an experienced Director of DevOps, Site Reliability & Infrastructure to lead teams responsible for the reliability, performance, security, and operational maturity of our SaaS platform. This leadership role involves ownership across DevOps, SRE, Infrastructure Operations, and Front-Line Technical Support, focusing on maintaining current environment stability while building the roadmap and practices for future scaling. The ideal candidate will balance strategy and execution, advise executives, lead through incidents, improve observability and processes, modernize infrastructure, and make strategic investment decisions to enhance reliability, scalability, measurability, cost-efficiency, and predictability.
Responsibilities
- Lead DevOps, SRE, Infrastructure, and Front-Line Support teams, establishing clear ownership, accountability, and operating standards.
- Own platform availability, performance, reliability, operational readiness, and service delivery.
- Lead incident management, root cause analysis, corrective actions, and post-incident reviews.
- Establish and continuously improve SLIs, SLOs, KPIs, service-level expectations, operational scorecards, and executive reporting.
- Create clear escalation paths across Support, SRE, DevOps, Engineering, Product, and other technical teams.
- Improve Tier 1/Tier 2 support effectiveness, including resolution rates, escalation quality, MTTR, and customer-impacting incidents.
- Build and maintain effective SOPs, runbooks, troubleshooting guides, and operational documentation.
- Own the operation, maintenance, and evolution of cloud, hybrid, and bare-metal infrastructure.
- Develop and execute a pragmatic future-state infrastructure roadmap aligned with business growth and technical requirements.
- Lead modernization initiatives to improve scalability, resilience, maintainability, and operational efficiency.
- Oversee high availability, disaster recovery, backup, business continuity, and capacity planning.
- Establish effective infrastructure governance balancing performance, security, reliability, cost, and complexity.
- Define and execute the monitoring and observability strategy across infrastructure, applications, platforms, and customer experience.
- Identify visibility gaps and improve alerting, issue detection, diagnosis, escalation, and resolution.
- Expand Infrastructure as Code, deployment automation, and operational automation to reduce manual work and improve consistency.
- Use metrics, retrospectives, and operational data to continuously improve reliability and team effectiveness.
- Evaluate emerging technologies, including AI-driven operational tools and AIOps, for performance or efficiency improvements.
- Own infrastructure and operational budgeting, forecasting, cost controls, and financial planning.
- Improve visibility into cloud, hosting, licensing, monitoring, support tooling, and other technology expenditures.
- Apply FinOps and cloud-governance principles to improve utilization and financial accountability.
- Identify and eliminate waste, overprovisioning, unused resources, and inefficient technology spending.
- Partner with Finance and Executive Leadership on future infrastructure investments and operating expenses.
- Build, mentor, and develop high-performing DevOps, SRE, Infrastructure, and Support teams.
- Establish clear roles, expectations, performance metrics, and development plans.
- Recruit and retain strong technical talent while building a culture of ownership, accountability, collaboration, and continuous improvement.
- Create clarity and momentum in a fast-changing environment.
- Build strong partnerships across Engineering, Product, QA, Security, Customer Success, Finance, and Executive Leadership.
- Communicate operational health, risks, priorities, investments, and progress clearly to technical and executive audiences.
Requirements
- 10+ years of experience across Infrastructure Operations, DevOps, SRE, Cloud Operations, or related disciplines.
- 5+ years of leadership experience managing engineering, DevOps, infrastructure, SRE, support, or service-delivery teams.
- Proven experience operating and supporting mission-critical SaaS platforms.
- Strong experience operating production cloud, hybrid, and/or bare-metal environments at scale.
- Deep understanding of Linux, networking, security, virtualization, storage, databases, infrastructure automation, and systems operations.
- Demonstrated experience improving monitoring, observability, incident response, operational maturity, and service reliability.
- Experience establishing effective SOPs, runbooks, operational standards, and support processes.
- Experience managing technology budgets, infrastructure costs, vendors, capacity, and operational expenditures.
- Strong planning, prioritization, decision-making, and organizational skills.
- Excellent communication skills with the ability to work effectively with engineers, business leaders, executives, partners, and customers.
- Experience with DevOps, SRE, or infrastructure transformation.
- Experience with Infrastructure as Code such as Terraform, Pulumi, Bicep, or CloudFormation.
- Experience with CI/CD and deployment automation.
- Experience with enterprise monitoring and observability platforms.
- Experience with Kubernetes, containerized workloads, and platform engineering.
- Experience with FinOps, cloud governance, and cost allocation.
- Experience with global SaaS platforms and 24x7 production environments.
- Experience with AI-driven operational tooling, automation platforms, or AIOps.
- Experience scaling technical service organizations through periods of business growth or transformation.
Skills
- DevOps
- Site Reliability Engineering (SRE)
- Infrastructure Operations
- Cloud Operations
- SaaS Platform Operations
- Linux
- Networking
- Security
- Virtualization
- Storage
- Databases
- Infrastructure Automation
- Systems Operations
- Monitoring
- Observability
- Incident Response
- Service Reliability
- SOPs
- Runbooks
- Support Processes
- Technology Budget Management
- Infrastructure Cost Management
- Vendor Management
- Capacity Planning
- Operational Expenditures
- Planning
- Prioritization
- Decision-Making
- Organizational Skills
- Communication Skills
- Infrastructure as Code (Terraform, Pulumi, Bicep, CloudFormation)
- CI/CD
- Deployment Automation
- Kubernetes
- Containerized Workloads
- Platform Engineering
- FinOps
- Cloud Governance
- Cost Allocation
- AI-driven Operational Tooling
- AIOps
Experience Level
- Director
- 10+ years of experience
- 5+ years of leadership experience
Education Level
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
About the Company
- We are a company focused on scaling our SaaS platform and evolving our technology organization for the future.
- We aim to strengthen reliability, modernize our platform, develop our teams, and create an operational foundation for continued growth.
