About the Role
The Senior AI Platform Operations Engineer is responsible for the reliability, operability, and controlled enablement of the organization’s AI platform, ensuring AI Platform services and solutions are production-ready, secure, observable, and compliant. This role enables the safe and scalable adoption of AI by ensuring solutions are deployed, monitored, supported, and continuously improved.
Responsibilities
- Administer and operate the AI platform to ensure availability, performance, and resilience.
- Monitor platform health using dashboards, logs, metrics, and alerts, and coordinate incident and service restoration.
- Lead operational triage, escalation coordination, and post-incident reviews.
- Track and report on service reliability indicators, incident trends, and operational performance.
- Enable approved AI use cases into production by ensuring environment readiness, dependency validation, operational readiness checklist completion, and structured service transition.
- Support platform lifecycle management through release coordination, change readiness validation, and maintenance and capacity planning.
- Ensure AI platform changes meet defined operational and control readiness criteria prior to release.
- Implement and maintain observability capabilities, including telemetry, logging, metrics, and traces.
- Analyze operational data to identify anomalies, recurring issues, and root-cause patterns.
- Implement AI Ops use cases such as alert correlation, anomaly detection, root-cause support, forecasting, and automation of repetitive operational tasks.
- Continuously improve operational efficiency through targeted automation and process optimization.
- Execute governance controls for AI solutions, including usage and access controls, data privacy, auditability, and human oversight.
- Ensure operational practices align with enterprise security policies, risk controls, and compliance requirements.
- Maintain documentation and evidence for audits, governance reviews, and control validation.
- Identify control gaps and escalate risks appropriately.
- Maintain operational visibility of AI platform assets for monitoring, support, and cost alignment.
- Validate asset ownership, relationships, and lifecycle status.
- Support ongoing audits to ensure AI assets and cost attribution remain accurate.
Requirements
- University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
- 5-7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations.
- Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.
- Experience with cloud platforms, observability, automation, configuration management, and integration patterns.
- Experience with Azure Automation runbooks (PowerShell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub-and-spoke), and App Service.
- Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
- Experience with CI/CD and automation tools such as Azure DevOps, GitHub Actions, and Logic Apps.
- Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies.
- Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka.
- Working knowledge of platform-supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB.
- Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.
- Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance controls, human-in-the-loop practices, and production monitoring.
- Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, and service reporting practices.
- Analytical and structured thinker with strong troubleshooting, root-cause analysis, prioritization, and continuous improvement skills.
- Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, and business teams.
- Experience creating technical documentation, operational procedures, support playbooks, dashboards, and user guidance materials.
- Knowledge of security, privacy, audit, and compliance considerations relevant to enterprise AI and platform operations.
Skills
- Platform Operations
- Site Reliability Engineering
- DevOps
- Cloud Operations
- Enterprise IT Operations
- Monitoring
- Incident Response
- Problem Management
- Service Restoration
- Operational Reporting
- Azure Automation
- PowerShell
- Python
- Azure AI
- Copilot
- AKS
- Virtual Networks
- App Service
- Azure Monitor
- Application Insights
- Grafana
- Azure DevOps
- GitHub Actions
- Logic Apps
- Bicep
- Terraform
- Azure Policy
- Key Vault
- API Management
- Service Bus
- Event Grid
- Apache Kafka
- Elastic
- Azure AI Search
- Cosmos DB
- DNA
- Fortinet
- Akamai
- AI/ML Operational Concepts
- Model Lifecycle Support
- ITIL/ITSM
- Troubleshooting
- Root-Cause Analysis
- Prioritization
- Continuous Improvement
- Technical Documentation
- Operational Procedures
- Support Playbooks
- Dashboards
- Security
- Privacy
- Audit
- Compliance
Experience Level
- Senior
Education Level
- University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
