Cloud Performance Engineering - Site Reliability Engineer at Smile Digital Health | CAN | Rezi

Cloud Performance Engineering - Site Reliability Engineer at Smile Digital Health

Cloud Performance Engineering - Site Reliability Engineer

Smile Digital Health · CAN

2 weeks ago

Cloud Performance Engineering - Site Reliability Engineer

Smile Digital Health · CAN

19 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now

About the Role

The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms. This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.

Responsibilities

  • Collaborate with Security Operations teams to define and implement best practices for Cloud Service Provider configuration (Azure and others).
  • Develop, implement, and coordinate a multi-tenant approach for service offerings (DB, Container platform, Authentication, Certificates, Product Registries).
  • Design and maintain performance testing strategies, frameworks, and environments in the cloud.
  • Develop and maintain cost/utilization tracking and attribution processes for all Cloud Service Providers.
  • Create documentation for Cloud Service Provider offerings, detailing use cases, best practices, and implementation.
  • Develop and maintain technical relationships with core Cloud Service Providers.
  • Implement and maintain a secure and scalable infrastructure platform for delivering Cloud Services applications.
  • Ensure internal and external SLAs meet and exceed expectations, and continuously monitor and improve system-centric KPIs.
  • Create tools for automating deployment, monitoring, and operations of the overall platform.
  • Participate in an on-call rotation for application support, incident management, and troubleshooting.
  • Provide ongoing maintenance and support of internal tools, improving system health and reliability.
  • Assist customers with on-site deployments as needed.
  • Implement and manage observability tools (logging, metrics, tracing) for performance insights, with a preference for Otel and Grafana Stack.

Requirements

  • Demonstrated expertise in cloud service providers and best practices for implementation and configuration, preferably managing Azure for multiple teams in a SaaS environment.
  • Proven experience with microservices architecture, with a strong focus on Java-based services.
  • Experience applying chaos engineering practices to evaluate and enhance system resiliency.
  • Skilled in troubleshooting performance issues, including analyzing time consumption, allocating resources, and recommending optimizations.
  • Familiar with performance testing methodologies and tools to assess system behavior under load.
  • Experience with deployment and usage of observability tools like Prometheus and the Grafana suite.
  • Proven experience designing and executing performance test plans (load, stress, soak, spike testing) to validate application services sustain 500+ transactions per second (TPS) within defined latency and error-rate thresholds.
  • Hands-on experience with performance/load testing tools such as JMeter, Gatling, Azure Load Testing.
  • Experience tuning and validating autoscaling (Kubernetes/OpenShift HPA, Azure scale sets) to meet required TPS under variable load without breaching cost or resource constraints.
  • Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing components to sustain target transaction rates.
  • Experience with Azure-native monitoring and diagnostics (Azure Monitor, Application Insights, Log Analytics) to correlate throughput, latency, and error metrics during test execution.
  • Proven experience with Security and Compliance (SOC2, HIPAA, ISO27001) best practices and implementing controls for high-velocity software delivery teams.
  • Proficiency in Terraform, Ansible, or Chef.
  • Expertise in troubleshooting, support escalation, on-call process optimization, and documenting knowledge.
  • Passionate about Infrastructure as Code, automation, and developing solutions that enable developers to move quickly and safely.
  • Familiarity with infrastructure management and operations lifecycle concepts and ecosystem.
  • Experience operating and maintaining production systems in a Linux and public cloud environment.
  • Prior experience working in high-performance or distributed systems.
  • Working knowledge of industry best practices regarding information security.
  • Previous experience building or maintaining a large-scale Cloud service.
  • Proven ability to prioritize and track multiple projects in parallel.

Skills

  • Cloud Service Provider configuration
  • Azure
  • Multi-tenant architecture
  • Service offerings
  • DB
  • Container platform
  • Authentication
  • Certificates
  • Product Registries
  • Performance testing
  • Cloud environments
  • Cost/utilization tracking
  • Technical relationship management
  • Infrastructure platform
  • SLA management
  • KPI monitoring
  • Automation tools
  • Deployment automation
  • Monitoring automation
  • Operations automation
  • On-call rotation
  • Incident management
  • Troubleshooting
  • System health
  • System reliability
  • On-site deployments
  • Observability tools
  • Logging
  • Metrics
  • Tracing
  • Otel
  • Grafana Stack
  • Microservices architecture
  • Java
  • Chaos engineering
  • System resiliency
  • Performance issue troubleshooting
  • Time consumption analysis
  • Resource allocation
  • System optimization
  • Prometheus
  • Grafana
  • Performance test plans
  • Load testing
  • Stress testing
  • Soak testing
  • Spike testing
  • JMeter
  • Gatling
  • Azure Load Testing
  • Autoscaling
  • Kubernetes HPA
  • OpenShift HPA
  • Azure scale sets
  • Kafka tuning
  • Messaging components
  • Queueing components
  • Azure Monitor
  • Application Insights
  • Log Analytics
  • Security and Compliance
  • SOC2
  • HIPAA
  • ISO27001
  • Terraform
  • Ansible
  • Chef
  • Support escalation
  • On-call process optimization
  • Knowledge documentation
  • Infrastructure as Code
  • Infrastructure management
  • Operations lifecycle
  • Linux
  • Public cloud
  • High-performance systems
  • Distributed systems
  • Information security

Location

  • Remote

Work Type

  • Remote Work Environment

Experience Level

  • New role
  • Prior experience working in high-performance or distributed systems
  • Previous experience building or maintaining a large-scale Cloud service

Salary/Compensations

  • $110,000 - $125,000 a year

Benefits

  • Remote Work Environment
  • Flexible Time Away From Work Policy including PTO, Personal and Sick Days
  • Competitive Salary and Health/Medical Benefits
  • RRSP/TFSA/401K Employee Contribution
  • Life and Disability
  • Employee Assistance Program
  • FHIR Study Program and Skillsoft Learning
  • Super HAPI Fun Club

About the Company

  • Smile Digital Health supports the mandate for #BetterGlobalHealth through its innovative health data platform and data management solutions, used in over 20 countries.
  • The company was ranked #19 on Deloitte's Technology Fast 50 Ranking for 2024.
  • Smile Digital Health's FHIR-based data liberation platform enables healthcare stakeholders to collect and exchange data.
  • The Smile platform helps organizations better manage healthcare data, generating and liberating structured healthcare data for effective delivery across care teams and health systems.
  • Smile discloses that artificial intelligence (AI) may be used in portions of the recruitment and selection process, such as resume screening or application assessment. All hiring decisions are ultimately made by qualified human decision-makers, and AI tools are used to support — not replace — fair and equitable hiring practices.
  • This position is a new role, created to support Smile’s continued growth and commitment to operational excellence.
  • Smile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success.
  • The company is dedicated to fostering a workplace that values diversity, equity, and inclusion.
  • Smile is big on creating a sense of belonging and empowering each other to bring their authentic selves to work.

Equal Opportunity

  • We welcome and encourage candidates of all backgrounds to apply.
  • Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.