Site Reliable Engineer (Canada - Remote) at AXON-Networks | CA | Rezi

Site Reliable Engineer (Canada - Remote) at AXON-Networks

Site Reliable Engineer (Canada - Remote)

AXON-Networks · CA

3 weeks ago

Site Reliable Engineer (Canada - Remote)

AXON-Networks · CA

23 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Site Reliable Engineer (Canada - Remote) role.

Rezi rewrites your resume against AXON-Networks's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Site Reliable Engineer (Canada - Remote) posting at AXON-Networks — free, in seconds.

About the Role

The Site Reliability Engineer will improve the availability, performance, scalability, and recoverability of AXON Networks cloud solutions by combining software engineering with hands-on NOC operations to make the complete cloud-to-device service path observable, supportable, and resilient at fleet scale. You will help establish practical SRE capabilities inside the NOC while partnering closely with Support, Operations, cloud, and DevOps Engineering, and participate in a sustainable on-call rotation to improve the NOC’s ability to diagnose customer-impacting issues.

Responsibilities

  • Own reliability outcomes for assigned cloud services.
  • Improve observability, capacity, resilience, and recovery.
  • Define and operationalize service-level indicators, service-level objectives, and actionable alerting.
  • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations, and routine production changes.
  • Lead technically during incidents, drive evidence-based learning, and ensure high-value corrective actions are completed.
  • Establish reliability baselines, SLIs, SLOs, and error budgets for cloud services and critical device-management workflows.
  • Trace failures across the end-to-end service path.
  • Identify fleet-wide and customer-specific failure patterns.
  • Contribute operability requirements and production evidence during design and readiness reviews.
  • Maintain NOC dashboards for service health, device reachability, provisioning success, command and telemetry performance, firmware adoption, and customer impact.
  • Participate in the NOC production on-call rotation and serve as a technical incident lead or senior troubleshooter.
  • Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging platforms, device-management sessions, and CPE behavior.
  • Coordinate evidence gathering and technical escalation with service-provider customers, Engineering, firmware, DevOps, and vendors.
  • Lead or contribute to post-incident reviews and convert recurring failures into prioritized and measurable corrective actions.
  • Develop production-grade software, scripts, and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance, and recovery.
  • Improve CI/CD and GitOps practices for operational software and infrastructure.
  • Manage or contribute to infrastructure as code, configuration as code, and reusable self-service patterns.
  • Measure NOC toil and partner with Automation & Tools Engineers to prioritize durable platform capabilities.
  • Develop capacity models for service-provider growth, managed-device populations, telemetry volume, messaging throughput, API demand, and rollout events.
  • Create and maintain runbooks, troubleshooting decision trees, service maps, device and cloud dependency records, known-error guidance, and operational knowledge.
  • Coach NOC and Support personnel on diagnosis, safe mitigation, evidence capture, and escalation.
  • Build self-service diagnostic views and tools that help the NOC determine scope, affected customers, device cohorts, likely fault domain, and next action.
  • Share reliability insights with Engineering and Product and contribute to reliability reviews, operational-readiness reviews, and continuous-improvement priorities.

Requirements

  • 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role.
  • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code.
  • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers.
  • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments.
  • Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code.
  • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis.
  • Strong troubleshooting & debugging skills in Kubernetes platforms.
  • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring.
  • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems.
  • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices.
  • Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.
  • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning.
  • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting.
  • Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams.
  • Bachelor’s degree in computer science, engineering or equivalent practical experience.
  • Experience supporting cloud-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices in a service-provider environment.
  • Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management.
  • Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes.
  • Understanding of access technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks interact.
  • Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices.
  • Experience building auto-remediation, safe self-service operations or internal reliability platforms.
  • Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment.

Skills

  • Python
  • Go
  • Java
  • Bash
  • Google Cloud Platform
  • Oracle Cloud Infrastructure
  • Kubernetes
  • Terraform
  • Helm
  • Git-based CI/CD
  • Policy-as-code
  • Linux
  • Containers
  • Prometheus
  • Grafana
  • OpenTelemetry
  • Apache Pulsar
  • TCP/IP
  • DNS
  • DHCP
  • TLS
  • Routing
  • NAT
  • Load balancing
  • TR-069/CWMP
  • TR-369/USP
  • TR-181
  • ACS
  • USP
  • Apache Kafka

Location

  • Irvine, CA USA

Work Type

  • Contract

Experience Level

  • 5+ years of experience

Education Level

  • Bachelor’s degree in computer science, engineering or equivalent practical experience

Salary/Compensations

  • 80 CAD/hr - 110 CAD/hr

About the Company

  • AXON Networks delivers a robust AI-driven, analytics-based orchestration platform and a wide portfolio of next-gen high-speed routers that leverage the newest Wi-Fi technologies. Together, these technologies give ISPs the ability to manage and troubleshoot their networks in real time, and to deliver an outstanding customer experience.
  • AXON Networks is a headquartered in Irvine, CA USA with Asia HQ in Singapore and also operating in Denmark, Spain and Vietnam.
  • AXON Networks is a trusted strategic partner for its customers, helping them evaluate their current technologies and business models, and creating and executing strategies that enable them to innovate faster, accelerate their digital transformations, and strengthen their relationships with consumers.

Equal Opportunity

  • AXON Networks promote equal opportunities in all our recruitment processes, ensuring non-discrimination on the basis of gender, age, origin, disability, or any other personal circumstances. We assess talent based on objective criteria and foster an inclusive and diverse working environment.