Talent.com
Royal Bank of Canada>
Site Reliability Engineer (SRE), Cloud OperationsRoyal Bank of Canada> • TORONTO, Canada
Site Reliability Engineer (SRE), Cloud Operations

Site Reliability Engineer (SRE), Cloud Operations

Royal Bank of Canada> • TORONTO, Canada
3 days ago
Job type
  • Full-time
Job description

Job Description

What is the opportunity?

Join the Platform Engineering & AI Operations team within OTK0, where you'll sit at the intersection of Site Reliability Engineering and intelligent infrastructure operations. This role offers the chance to shape how the bank operates, monitors, and self-heals its private and public cloud platforms — from OpenShift clusters and Kafka environments to self-healing automation systems. You'll work on real problems at enterprise scale: reducing toil for NOC, Data Center, and Branch teams, building automation that eliminates manual work, and establishing reliable operational practices. If you want to move beyond traditional ops into the future of intelligent, autonomous infrastructure operations — this is the role.

What will you do?

  • Support highly scalable, secure, and highly available architectures across private and public cloud platforms (Kubernetes/OpenShift, ECE, Confluent Kafka).
  • Write code and scripts to automate infrastructure workflows and eliminate toil, including automation pipelines that reduce manual intervention across Data Center, Branch, and NOC operations.
  • Extend self-healing automation capabilities built on Ansible, automating routine operational tasks (e.g., CPU remediation) to reduce manual intervention.
  • Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points, ensuring alignment with security, reliability, and regulatory requirements.
  • Collaborate with platform teams to provide technical feedback, contribute code changes to shared repositories, and establish data standards and pipelines (e.g., ServiceNow, Prometheus) that support operational excellence.
  • Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform for deployment validation and self-healing remediation workflows.
  • Minimize risk of reliability failures related to durability, availability, performance, and correctness, leveraging proactive alerting and anomaly detection.
  • Participate in on-call rotation for platform support, incident management, and troubleshooting, triaging incidents via Grafana, Prometheus, Dynatrace, and PagerDuty.

What do you need to succeed?

Must-have

  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
  • Strong working knowledge of Kubernetes/OpenShift administration and troubleshooting in enterprise environments.
  • Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Python scripting (core to infrastructure automation and platform development).
  • Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
  • Experience with incident management processes, on-call rotations, and post-incident review practices.
  • Familiarity with capacity planning, threshold-based alerting, and performance trend analysis.
  • Understanding of security and compliance fundamentals, including vulnerability assessment and remediation tracking.
  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).

Nice-to-have

  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
  • Hands-on experience with public cloud platforms (AWS, Azure, GCP) in hybrid or multi-cloud environments.
  • Experience with GPU/compute infrastructure for ML inference workloads

What’s in it for you?

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.

  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • Flexible work/life balance options
  • Opportunities to do challenging work
  • Opportunities to take on progressively greater accountabilities
  • Access to a variety of job opportunities across business

#LI-post

#TECHPJ

Job Skills

Agile Methodology, Ansible Tower, Group Problem Solving, IT System Administration, IT Systems Integration, Kubernetes, Linux, Organizational Leadership, Product Services, RedHat OpenShift Administration, Red Hat OS Administration, Software Development Life Cycle (SDLC), System Applications, System Integration Testing (SIT), Systems Software

Additional Job Details

Address:

RBC CENTRE, 155 WELLINGTON ST W:TORONTO

City:

Toronto

Country:

Canada

Work hours/week:

37.5

Employment Type:

Full time

Platform:

TECHNOLOGY AND OPERATIONS

Job Type:

Regular

Pay Type:

Salaried

Posted Date:

2026-07-31

Application Deadline:

2026-08-28

Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above

Our Employment Opportunities

At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.

Join our Talent Community

Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.

Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com.

RBC is presently inviting candidates to apply for this existing vacancy. Applying to this posting allows you to express your interest in this current career opportunity at RBC. Qualified applicants may be contacted to review their resume in more detail.

Create a job alert for this search

Site Reliability Engineer (SRE), Cloud Operations • TORONTO, Canada

Similar jobs

Senior Site Reliability Engineer

Guidewire Softwaretoronto, on, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, ON, CA
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Lead Site Reliability Engineer (Sre) - C$75,900 - C$141,900 Par An

BMO Groupe financierToronto, Canada
Part-time

Date limite pour présenter sa candidature : 05/30/2026Adresse : 4100 Gordon Baker RoadGroupe de famille d'emploi : TechnologieHybrid role (2 days/week in Scarborough office).Out of province can... Show more

 • Promoted

Senior SRE: Global SaaS Platform, Kubernetes & Cloud

Kong Inc.Toronto, ON, CA
Full-time

A leading cloud API technology developer is seeking a Senior Site Reliability Engineer to enhance and operate their global SaaS platform.This hands-on role involves designing and automating product... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer

PheedLoopToronto, ON, CA
Full-time

Build the tech behind live events.PheedLoop's mission is to help organizers turn ordinary events into unforgettable experiences with event technology that is bold, intuitive, and built to bring peo... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Site Reliability Engineer

KyndrylToronto, ON, CA
Full-time +1

Join to apply for the Site Reliability Engineer role at Kyndryl.Direct message the job poster from Kyndryl.Recruitment & Strategic Staffing @Kyndryl | Partnering with IT Consultants in Financial Se... Show more

 • Promoted

Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc.Toronto, ON, CA
Full-time

A leading supply chain solutions provider is seeking a Site Reliability Engineer to optimize and ensure the reliability of their cloud infrastructure across AWS and Kubernetes.This role emphasizes ... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsToronto, ON, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Lead Site Reliability Engineer for Multi-Cloud

MongoDBToronto, ON, CA
Full-time

Lead the charge in securing multi-cloud communications as a skilled Site Reliability Engineer.Bring your network expertise to develop and maintain critical infrastructure for service connectivity.T... Show more

 • Promoted

Site Reliability Engineer

Future Secure AIToronto
Full-time

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Show more

 • Promoted

Site Reliability Engineer for Cloud Infrastructure Management

NewtonToronto, ON, CA
Full-time

Be a pivotal Site Reliability Engineer focused on improving infrastructure resilience and reliability.Collaborate remotely to drive operational success and enhance system performance in a dynamic e... Show more

 • Promoted

Site Reliability Engineer

Momentum Financial Services GroupToronto, ON, CA
Full-time

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Show more

 • Promoted

Site Reliability Engineer (Senior or Staff), Deployments

AlleyCorpToronto, ON, CA
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto, ON, CA
Permanent

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificToronto, ON, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Site Reliability / DevOps Engineer

Infotek Consulting Inc.Toronto, ON, CA
Full-time

This range is provided by Infotek Consulting Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.CAN – Site Reliability / DevOps Engineer (Exper... Show more