About the Role
Build and maintain the monitoring, observability, and infrastructure automation that keeps our cloud platform reliable, working within a VMware Cloud Foundation (VCF) and Kubernetes environment.
Responsibilities
- Deploy, configure, and maintain Zabbix for system and network monitoring.
- Build and maintain Prometheus exporters and Grafana dashboards for capacity planning, performance metrics, and operational visibility.
- Configure and manage Opsgenie for alert routing, escalation policies, and on-call schedules.
- Analyze alert noise, tune thresholds, and reduce false positives to keep alerting actionable.
- Integrate monitoring systems with ticketing, communication, and incident management tools.
- Maintain and optimize monitoring infrastructure, including database tuning, storage management, and high availability.
- Write and maintain Ansible playbooks for deploying and configuring monitoring infrastructure.
- Use Terraform for provisioning infrastructure resources.
- Build and maintain GitLab CI/CD pipelines for automated testing, linting, and deployment of monitoring and infrastructure code.
- Follow infrastructure-as-code practices, including version control and peer reviews.
- Act as a first responder for monitoring-related incidents and alerts.
- Investigate and resolve performance issues, outages, or anomalies detected by monitoring systems.
- Escalate issues to appropriate teams (network, security, infrastructure) when needed.
- Document incidents, root causes, and resolutions, and contribute to post-incident reviews.
- Provide technical support to internal teams on monitoring tools and dashboards.
- Work within VMware vCenter / VCF and Kubernetes environments to support monitoring and infrastructure needs.
- Manage notification infrastructure, including SMTP relay configuration and delivery troubleshooting.
- Support compliance requirements (ISO 27001, SOC 2) by maintaining audit logging, access controls, and security configurations for monitoring systems.
Requirements
- Diploma or degree in Computer Science, IT, or a related field, or equivalent practical experience.
- Eligible to obtain Secret Level Clearance within the first 3 months of employment.
- Canadian Citizen with 10+ years of verifiable police background history.
- Hands-on experience with Zabbix or comparable monitoring tools (Nagios, Icinga, Checkmk).
- Working knowledge of Prometheus and Grafana, including writing exporters, building dashboards, and PromQL.
- Experience with alert management and on-call tooling (Opsgenie, PagerDuty, or similar).
- Comfort with Linux systems administration.
- Proficiency in scripting and automation using Bash and Python.
- Experience with at least one IaC tool (Ansible, Terraform).
- Familiarity with CI/CD pipelines (GitLab CI, GitHub Actions, Jenkins, or similar).
- Understanding of networking fundamentals (TCP/IP, DNS, SNMP, bandwidth/latency concepts).
- Basic database administration (PostgreSQL or MySQL) for monitoring tool backends.
- Strong diagnostic and troubleshooting skills.
- Clear written and verbal communication skills.
- Attention to detail.
- Comfort working independently in a remote environment while collaborating across teams.
- Ability to stay composed during incidents and work under time pressure.
Skills
- Zabbix
- Prometheus
- Grafana
- Opsgenie
- Ansible
- Terraform
- GitLab CI/CD
- VMware Cloud Foundation (VCF)
- Kubernetes
- Bash
- Python
- Linux
- TCP/IP
- DNS
- SNMP
- PostgreSQL
- MySQL
Location
- Canada
- Remote
Work Type
- Remote-first
Education Level
- Diploma or degree in Computer Science, IT, or a related field, or equivalent practical experience.
Benefits
- Competitive compensation package
- Flexible time-off for vacation, plus illness & personal days
- Comprehensive health and dental benefits
- GRSP & 401k Matching Program
About the Company
- ThinkOn is a managed infrastructure services provider (MISP) with a global data center footprint, focused on empowering partners to do whatever they need to do with their data.
- ThinkOn’s team of data-obsessed experts protect clients’ data like it’s their own, making it more resilient, secure, actionable, and searchable.
- ThinkOn is Channel First and works with a global network of value-add resellers and managed service providers to provide creative, turnkey Infrastructure-as-a-Service (IaaS), Disaster Recovery-as-a-Service (DRaaS), and Backup-as-a-Service (BaaS) solutions and data management services that are fast, flexible, scalable, highly secure, and cost-effective with predictable pricing and no hidden fees.
- ThinkOn has data centers located across North America, the United Kingdom, Australia, and the Caribbean.
- Recognized for its substantial growth and global success, ThinkOn has been named to several notable lists, including The Globe and Mail’s Report on Business ranking of “Canada’s Top Growing Companies” (2021 and 2022), the Deloitte Technology Fast 500 (2021 and 2022), the Deloitte Technology Fast 50 (2021), CIOReview’s “Most Promising Backup Solution Provider” (2021), the Canadian Business “Growth List,” and Channel Daily News’ “Top 100 Solution Providers” (2021).
- ThinkOn is headquartered in Toronto, Ontario.
Equal Opportunity
- ThinkOn is committed to a workforce that is reflective of diverse populations.
- We welcome applications from qualified individuals from all backgrounds.
- In accordance with the Accessibility for Ontarians with Disabilities Act (AODA) and accessibility standards across Canada, ThinkOn provides accommodations to job applicants with disabilities throughout the recruitment process.
- If you require accommodations, please let us know and we will work with you to meet your needs.
- We are committed to a selection process and work environment that is inclusive, equitable, accessible, and adheres to our corporate values.
