Talent.com
Smile Digital Health
Cloud Performance Engineering Site Reliability Engineer ( Remote Canada)Smile Digital Health • Toronto, Ontario, Canada
Cloud Performance Engineering Site Reliability Engineer ( Remote Canada)

Cloud Performance Engineering Site Reliability Engineer ( Remote Canada)

Smile Digital Health • Toronto, Ontario, Canada
16 days ago
Salary
CA$110,000.00 yearly
Job type
  • Full-time
  • Remote
Job description
Working for a company like Smile Digital Health means supporting our mandate for #BetterGlobalHealth. We strive towards this goal every day and the results can be seen in the impact of our innovative health data platform and data management solutions which are used in over 20 countries. We were #19 on Deloittes Technology Fast 50 Ranking for 2024!Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.At its heart the Smile platform enables people and organizations to better manage healthcare data. We helpgenerate andliberate structured healthcare data to ensure effective delivery across care teams and health systems bringing #BetterGlobalHealth to patients everyday!
Apply today and find plenty of reasons to SMILE!

The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability scalability and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health its clients and partners.

This role designs and automates performance testing frameworks integrates them into CI/CD pipelines and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering product and security teams the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.

Responsibilities:

  • Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
  • Develop implement and coordinate a multi-tenant approach around service offerings for databases container platforms authentication certificates and product registries.
  • Design develop and maintain cloud performance testing strategies frameworks and environments to validate application scalability reliability and resiliency.
  • Develop and automate load stress spike and endurance (soak) testing as part of CI/CD pipelines.
  • Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
  • Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
  • Create documentation detailing Cloud Service Provider offerings implementation patterns and best practices.
  • Develop and maintain technical relationships with our core Cloud Service Providers.
  • Implement and maintain secure scalable infrastructure platforms for delivering cloud services.
  • Ensure internal and external SLAs are consistently met or exceeded while continuously monitoring and improving system performance reliability and availability.
  • Create tools for automating deployment monitoring and platform operations.
  • Implement and manage observability solutions (logging metrics tracing) using OpenTelemetry Prometheus Grafana Azure Monitor and related technologies to provide actionable performance insights.
  • Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.

Requirements:

  • 5 years of experience with Cloud Service Providers and best practices around implementation and configuration preferably managing Azure environments supporting SaaS products.
  • Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
  • Strong experience in Cloud Performance Engineering including performance analysis capacity planning scalability testing and optimization of distributed cloud-native applications.
  • Proven experience working with microservices architecture with a strong focus on Java-based services.
  • Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
  • Strong experience designing and executing performance testing strategies including load stress spike and endurance (soak) testing to validate application scalability and defined latency and error-rate thresholds.
  • Hands-on experience with performance testing tools such as JMeter Gatling Azure Load Testing or k6.
  • Experience validating application services sustaining 500 transactions per second (TPS) while meeting defined performance objectives.
  • Hands-on experience deploying and managing containerized applications using Docker and Kubernetes including autoscaling and performance optimization.
  • Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (IaC).
  • Experience tuning Kafka (partitioning consumer group sizing throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
  • Hands-on experience implementing and using observability platforms including OpenTelemetry Prometheus Grafana Azure Monitor Application Insights and Log Analytics.
  • Proven experience with Security and Compliance (SOC 2 HIPAA ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.
$110000 - $125000 a yearSmile discloses that artificial intelligence (AI) may be used in portions of the recruitment and selection process such as resume screening or application assessment. All hiring decisions are ultimately made by qualified human decision-makers and AI tools are used to support not replace fair and equitable hiring practices.This position is a replacement role created to support Smiles continued growth and commitment to operational excellence.
Some of the benefits we offer:* Remote Work Environment* Flexible Time Away From Work Policy including PTO Personal and Sick Days* Competitive Salary and Health/Medical Benefits* RRSP/TFSA/401K Employee Contribution* Life and Disability* Employee Assistance Program* FHIR Study Program and Skillsoft Learning* Super HAPI Fun Club
Smiles core values include respect inclusion embracing our differences and celebrating shared values because our people are the foundation of our success. We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work. We are dedicated to fostering a workplace that values diversity equity and inclusion.We welcome and encourage candidates of all backgrounds to apply. Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.We may use artificial intelligence (AI) tools to support parts of the hiring process such as reviewing applications analyzing resumes or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed please contact us.

Required Experience:

IC


Employment Type : Full-Time
Experience: years
Vacancy: 1
Yearly Salary Salary: 110000 - 125000
Create a job alert for this search

Cloud Performance Engineering Site Reliability Engineer ( Remote Canada) • Toronto, Ontario, Canada

Similar jobs

Staff Site Reliability Engineer - Confluent Incident Management & Reliability

IBMToronto
Full-time

Your Role and Responsibilities.Confluent Cloud processes millions of events per second across AWS, GCP, and Azure.When incidents happen in a multi‑cloud streaming platform, they happen at scale—dat... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Site Reliability Engineer

PheedLoopToronto, ON, CA
Full-time

Build the tech behind live events.PheedLoop's mission is to help organizers turn ordinary events into unforgettable experiences with event technology that is bold, intuitive, and built to bring peo... Show more

 • Promoted

Senior Platform Engineer - Remote, Scale & Reliability

Lillio (formerly HiMama)Toronto, ON, CA
Remote
Full-time

A leading EdTech company in Canada is seeking a Senior Platform Engineer to enhance system reliability and performance while contributing to impactful software solutions.The role involves making te... Show more

 • Promoted

Remote Platform Engineer — Cloud & Kubernetes Ops

PlanetToronto, ON, CA
Remote
Full-time

A leading global space and data company is seeking a Software Engineer in Platform Operations.This full-time remote role prioritizes building and operating cloud infrastructure supporting engineeri... Show more

 • Promoted

Senior Site Reliability Engineer

Guidewire SoftwareToronto, Ontario, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Site Reliability Engineer For Cloud Infrastructure Management

NewtonToronto, Canada
Full-time

Be a pivotal Site Reliability Engineer focused on improving infrastructure resilience and reliability.Collaborate remotely to drive operational success and enhance system performance in a dynamic e... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Senior Site Reliability Engineer - $111,100 - $166,700 A Year - Remote

ThinkificNorth York, Canada
Remote
Full-time

Senior Site Reliability Engineer to optimize platform infrastructure, focusing on scaling, security, and performance through cloud-native practices, Kubernetes, and AWS. Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsToronto, ON, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Cloud Reliability Engineer - Aws & Observability - C$107,000 - C$157,300 A Year

A leading software companyToronto County, Canada
Full-time

Seeking a Senior Cloud Reliability Engineer to enhance API services' reliability and scalability on AWS.This hybrid role focuses on optimizing production workloads and CI/CD. Show more

 • Promoted

Platform Engineer - Cloud Infra, CI/CD & SRE (Remote)

Fiat RepublicToronto
Remote
Full-time

A leading fintech company is seeking a Platform Engineer in Toronto, Canada.This role focuses on managing scalable infrastructure using Google Cloud Platform services and developing CI/CD pipelines... Show more

 • Promoted

Senior Site Reliability Engineer, Platform & Cloud Finops - C$200,000 - C$300,000 A Year

HopperEast York, Canada
Full-time

Seeking a Senior Site Reliability Engineer for Hopper's Cloud FinOps team.Responsibilities include optimizing infrastructure costs, ensuring scalability, reliability, and security, and particip... Show more

 • Promoted

Senior Site Reliability Engineer I - C$183,000 - C$203,000 A Year - Remote

InstacartNorth York, Canada
Remote
Full-time

Senior Site Reliability Engineer to maintain platform operations, optimize performance, and develop scalable infrastructure.Will lead incident management, monitor systems, and deploy automation tools. Show more

 • Promoted

Senior Site Reliability Engineer - C$140,000 - C$155,000 A Year - Remote

McGraw HillToronto County, Canada
Remote
Full-time

Seeking a Senior Site Reliability Engineer to build and support reliable, high-capacity systems for learning platforms.Responsibilities include automating cloud infrastructure, optimizing performan... Show more

 • Promoted

Senior Cloud Reliability Engineer - $157,300 A Year

Autodesk, Inc.Toronto County, Canada
Full-time

This role focuses on ensuring cloud application reliability and performance, designing infrastructure, and optimizing AWS workloads. Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Site Reliability Engineer

Momentum Financial Services GroupToronto
Full-time

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Show more

 • Promoted

Senior Site Reliability Engineer Ii - Remote, Scale-Focused - C$183,000 - C$203,000 A Year - Remote

Leading Grocery Delivery ServiceNorth York, Canada
Remote
Full-time

Seeking a Senior Site Reliability Engineer to ensure platform performance, establish incident management, and oversee scalable infrastructure strategies.Requires programming and incident management... Show more