Site Reliability Engineer at Boson AI | Toronto | Rezi

Site Reliability Engineer at Boson AI

Site Reliability Engineer

Boson AI · Toronto

2 weeks ago

Site Reliability Engineer

Boson AI · Toronto

19 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

Boson AI builds production-grade AI systems for natural, capable, and useful AI communication. This role focuses on building and operating the infrastructure for large-scale AI training and serving, including networks, GPU clusters, storage, scheduling, and operational tooling. The ideal candidate has deep strength in at least one area (networking, cluster scheduling, storage, GPU systems, or AI infrastructure) and the curiosity to collaborate across others.

Responsibilities

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Own and automate operational workflows across networking, compute allocation, storage, GPU/server configuration, or AI platforms.
  • Build monitoring, alerting, runbooks, and incident-response practices.
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.
  • Improve provisioning, configuration management, testing, and deployment automation.
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management.
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards.

Requirements

  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role.
  • Strong hands-on expertise in at least one of the following: Networking (firewalls, switching, routing, ASN/BGP configuration, InfiniBand), Cluster and systems allocation (Kubernetes, SLURM, MAAS), Distributed storage (Ceph), GPU and server administration (CUDA drivers, firmware, BIOS, hardware troubleshooting), or AI training/model-serving infrastructure.
  • Experience operating production systems with a focus on availability, performance, security, and automation.
  • Strong Linux administration and scripting skills.
  • A systematic approach to troubleshooting across multiple layers of a complex system.
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team.

Skills

  • Networking
  • Cluster scheduling
  • Storage
  • GPU systems
  • AI infrastructure
  • Linux administration
  • Scripting
  • Kubernetes
  • SLURM
  • MAAS
  • Ceph
  • CUDA
  • NVIDIA GPUs
  • NCCL
  • InfiniBand
  • RDMA
  • RoCE
  • Ethernet
  • Terraform
  • Ansible
  • Prometheus
  • Grafana
  • AWS
  • GCP
  • Azure

Location

  • Toronto
  • Remote

Work Type

  • Remote

Experience Level

  • 4+ years

Salary/Compensations

  • $125,000 - $250,000

About the Company

  • Boson AI is building AI systems for real-world, business-critical use.
  • If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we’d love to hear from you.