About the Role
Wolters Kluwer is seeking a Lead AppOps Engineer to ensure the stability, reliability, and operational excellence of enterprise applications. This role focuses on managing the end-to-end operational lifecycle, application runtime health, incident management, and release readiness within a mature engineering environment.
Responsibilities
- Own production reliability for critical applications; define and track SLOs, error budgets, and capacity/performance baselines.
- Lead major incident response, drive clear business/technical communications, and ensure data-driven root cause analysis with preventative actions.
- Direct release and change operations: assess risk, enforce readiness gates, validate post-deployment health, and improve change success rate.
- Architect operational observability: design dashboards, alert strategies, log/trace pipelines, and runbook automation for rapid diagnosis and recovery.
- Establish and continuously improve operational standards, guardrails, and runbooks; automate repetitive tasks to reduce toil.
- Partner with engineering on resiliency patterns (circuit breakers, bulkheads, graceful degradation, retries) and performance tuning.
- Plan and execute capacity management, scaling strategies, and DR/BCP readiness, including failover testing and scenario exercises.
- Champion security-by-default in operations: secrets hygiene, patch/vulnerability remediation, certificate/DNS management, least-privilege access.
- Mentor AppOps engineers; provide technical guidance, code/review for automation, and develop on-call excellence.
- Drive service reviews with stakeholders; publish operational KPIs (MTTR, change success rate, incident rate) and lead continuous improvement roadmaps.
- Monitor application health, availability, and performance across environments; proactively identify issues and optimize application behavior.
- Triage, investigate, and resolve production incidents; participate in root cause analysis and drive long-term fixes.
- Coordinate and execute application deployments, ensure release readiness, validate post-deployment health, and collaborate with engineering teams for smooth rollouts.
- Maintain application environments, configuration baselines, secrets, access controls, and platform dependencies, ensuring consistency and compliance.
- Implement and maintain dashboards, alerts, and log pipelines using enterprise observability tools to ensure system transparency and rapid diagnosis.
- Develop and enhance runbooks, automate repeatable workflows, reduce manual toil, and improve operational efficiency.
- Track key reliability metrics, enforce operational standards, and drive continuous optimization to meet or exceed service commitments.
- Support change reviews, evaluate operational risks, ensure compliance with WK change processes, and validate operational readiness for all changes.
- Ensure adherence to security standards, support vulnerability remediation efforts, and maintain compliance with organizational policies.
- Partner with Engineering, CloudOps, Security, Compliance, and other teams to resolve issues, improve service quality, and enhance application resilience.
- Perform other duties as assigned by management.
- On call rotation responsibilities with the Service Delivery and Operations Team
Requirements
- Advanced expertise in operating applications on Azure and/or AWS, including networking, load balancers, DNS, certificates, storage, and messaging services.
- Strong knowledge of application operations in cloud environments (Azure/AWS).
- Hands-on with observability stacks (Datadog, Grafana/Prometheus, ELK/OpenSearch, Open Telemetry) and alert engineering.
- Experience with incident management, RCA, and operational troubleshooting.
- Strong practical understanding of CI/CD concepts and collaboration with release teams; experience validating releases in lower/production environments.
- Familiarity with infrastructure components: load balancers, storage networking, DNS and certificates.
- Proficiency in automation and scripting (PowerShell, Bash, Python) to build runbooks, health checks, and remediation workflows.
- Experience with deployment strategies (blue/green, rolling, canary) and traffic management.
- Security and compliance in operations: vulnerability remediation, secrets and key management, audit readiness.
- Ability to interpret logs, metrics, traces, and performance data.
- Experience managing multi-environment application lifecycles (Dev, QA, UAT, Prod).
- Infrastructure as Code (IaC): Terraform (modules, workspaces), Azure ARM/Bicep or AWS CloudFormation; policy-as-code and environment drift detection.
- Ability to ensure 24x7 application reliability and operational excellence.
- Manage end-to-end application lifecycle including deployments, configurations, and environment health.
- Collaborate with engineering, CloudOps, and Security teams to ensure smooth operations.
- Own operational KPIs such as uptime, MTTR, change success rate, and SLA/SLO adherence.
- Perform release coordination, deployment validation, and post-release monitoring.
- Lead incident response, communication, and escalation handling.
- Participate in change management and risk assessments for all application changes.
- Maintain runbooks, SOPs, and operational documentation.
- Drive continuous improvement for operational workflows and process maturity.
- Support audit, compliance, and security requirements for applications.
- Proven experience leading incident response, conducting RCAs, and implementing preventative controls.
- Excellent communication, stakeholder management, and mentoring skills in global, fast-paced environments.
- Strong understanding of Software Engineering Principals.
- Industry recognized Kubernetes Certification.
Skills
- Application Operations
- Production Support
- Cloud-Platform Application Management
- DevSecOps
- CI/CD
- Application Runtime Health
- Operational Workflows
- Incident Management
- Release Readiness
- Environment Reliability
- Monitoring
- Performance Optimization
- Escalations
- Alerting
- Logging
- Observability
- Runbook Automation
- Resiliency Patterns
- Capacity Management
- Scaling Strategies
- Disaster Recovery (DR)
- Business Continuity Planning (BCP)
- Security Best Practices
- Vulnerability Remediation
- Secrets Management
- Certificate Management
- DNS Management
- Access Control
- Azure
- AWS
- Networking
- Load Balancers
- Storage Services
- Messaging Services
- Datadog
- Grafana
- Prometheus
- ELK Stack
- OpenSearch
- Open Telemetry
- Root Cause Analysis (RCA)
- Deployment Strategies
- Traffic Management
- Infrastructure as Code (IaC)
- Terraform
- Azure ARM
- Bicep
- AWS CloudFormation
- Policy as Code
- Environment Drift Detection
- Software Engineering Principles
- Kubernetes
Location
- Hybrid
Work Type
- Full-time
Experience Level
- 8-10 Years
Education Level
- Bachelor’s degree in computer science, Information Systems, or a related field.
- Vendor certifications preferred: Azure Administrator/Architect or AWS SysOps/DevOps Professional; ITIL Foundation (or higher).
- Terraform Associate/Professional (or equivalent IaC certification) preferred; SRE Foundation a plus.
Salary/Compensations
- $118,300.00 - $207,400.00 USD
Benefits
- Medical, Dental, & Vision Plans
- 401(k)
- FSA/HSA
- Commuter Benefits
- Tuition Assistance Plan
- Vacation and Sick Time
- Paid Parental Leave
About the Company
- Wolters Kluwer offers a wide variety of competitive benefits and programs to help meet your needs and balance your work and personal life.
Equal Opportunity
- Our Interview Practices: To maintain a fair and genuine hiring process, we kindly ask that all candidates participate in interviews without the assistance of AI tools or external prompts. Our interview process is designed to assess your individual skills, experiences, and communication style. We value authenticity and want to ensure we’re getting to know you—not a digital assistant. To help maintain this integrity, we ask to remove virtual backgrounds and include in-person interviews in our hiring process. Please note that use of AI-generated responses or third-party support during interviews will be grounds for disqualification from the recruitment process.
- Applicants may be required to appear onsite at a Wolters Kluwer office as part of the recruitment process.
