Talent.com
Aviva
Senior Manager, Site Reliability & Infrastructure EngineeringAviva • Markham, ON, Canada
Senior Manager, Site Reliability & Infrastructure Engineering

Senior Manager, Site Reliability & Infrastructure Engineering

Aviva • Markham, ON, Canada
2 days ago
Salary
CA$125,000.00 yearly
Job type
  • Full-time
Job description

Individually we are people, but together we are Aviva. Individually these are just words, but together they are our Values – Care, Commitment, Community, and Confidence.

At Aviva Canada, we put people first, our employees, our customers, and our communities. We’re proud of a culture built on care, inclusion, and collaboration, where your voice matters and your growth is supported. We’re not just about insurance; we’re about making a real difference by protecting what matters most.

The Senior Manager, Site Reliability & Infrastructure Engineering will lead Aviva Canada’s evolution toward an engineering-led reliability and resilience capability across critical applications, platforms and services. This role will work across on-premises, AWS, SaaS/vendor-hosted and hybrid application environments, partnering with Engineering Operations, Cloud & Platform Engineering, Application Engineering, Cybersecurity, Operational Resilience, Business Continuity, Enterprise Architecture, Vendor Management and key partners including AWS, DXC, CGI, Snowflake and other SaaS providers.

The focus is to shift Aviva from periodic recovery and resilience testing to proactive, measurable reliability engineering. This uses observability platforms like Dynatrace, incident response tools such as PagerDuty or equivalent, and modern recovery platforms like Rubrik to improve service availability, operational insight, recovery readiness, and customer outcomes.

What you’ll do

Site Reliability Engineering

  • Establish and mature SRE practices across critical applications and platforms, including SLIs, SLOs, SLAs, service health indicators, post-incident reviews and continuous reliability improvement.
  • Partner with application, platform, infrastructure and vendor teams to embed reliability requirements into design, development, operational acceptance, release and production support processes.
  • Drive improvements in incident response, problem management, root cause analysis, alert quality, critical issue workflows, service health reporting and reduction of repeat incidents.
  • Improve Infrastructure Ops via automation, self-healing patterns, runbook improvements, AI-assisted operations and repeatable engineering practices.
  • Use observability insights to improve infrastructure availability, performance, capacity planning, & operational readiness.

Technology Resilience, Backup & Cyber Recovery

  • Maintain and evolve Aviva Canada’s technology resilience capability across on-premises platforms, AWS, critical applications, integration services and SaaS/vendor-hosted platforms.
  • Ensure resilience planning validates end-to-end recoverability, including application, data, integration, identity, network, platform, cloud and vendor dependencies.
  • Support annual BCP, DR and technology resilience testing, including RTO/RPO validation, dependency mapping, recovery sequencing, test evidence, gap management and remediation tracking.
  • Own and mature backup and cyber recovery practices, including hands-on use of Rubrik or a comparable enterprise data protection platform for backup policy management, immutability, restore validation, coordinating recovery procedures, reporting, evidence capture and operational support for critical workloads.
  • Partner with Cybersecurity on ransomware and destructive cyber event readiness, including clean recovery, backup integrity validation, isolated recovery environments, cyber recovery vault operations and restoration of critical services.
  • Ensure resilience and recovery requirements are embedded into architecture, cloud migration, vendor onboarding, operational readiness, change delivery and service governance.

Leadership

  • Lead, manage and coach group of engineers working across reliability, observability, production support, technology resilience and recovery practices.
  • Build a clear operating model for SRE and technology resilience ownership across application teams, infrastructure teams, cloud/platform teams, cybersecurity, operational resilience and third-party providers.
  • Define reliability and resilience reporting for senior leaders, including service health, SLO performance, incident trends, MTTR, alert quality, recovery readiness, open risks and remediation status.
  • Partner with Operational Resilience, Business Continuity, Technology Risk and Cybersecurity teams to ensure evidence, controls and remediation plans are audit-ready and aligned to regulatory expectations.
  • Influence engineering and operations teams to adopt reliability-by-design, automation-first and evidence-based ways of working.
  • Serve as a designated point of contact for material risks relating to service reliability, monitoring gaps, application recoverability, backup coverage, cyber recovery readiness, RTO/RPO gaps and vendor resilience ambiguity.
  • Partner with key infrastructure managed service providers to continuously improve service quality, strengthen operational performance, hold vendors accountable for meeting agreed service levels, outcomes, remediation commitments and continuous improvement targets.

What you’ll bring

  • 10+ years of technology experience across Site Reliability Engineering, Infrastructure Engineering, Platform Engineering, Cloud, Technology resilience, Disaster recovery, Cyber recovery or related disciplines.
  • 5+ years of experience leading & managing technical teams, with the ability to mentor engineers, set direction, manage priorities and high visible projects.
  • Experience with enterprise observability tooling such as Dynatrace, Grafana, Datadog or equivalent platforms.
  • Hands-on experience operating Rubrik or a comparable enterprise backup and cyber recovery platform, including backup policy configuration, immutable backup concepts, restore testing, recovery workflow execution, access controls, reporting, evidence capture and support for critical workloads.
  • Experience supporting distributed, business-critical applications across hybrid environments, including on-premises infrastructure, cloud platforms, APIs, middleware, databases, containers and SaaS/vendor-hosted systems.
  • Working knowledge of AWS or comparable cloud platforms, including cloud resilience, infrastructure as code, monitoring, logging, backup/recovery, identity, network dependencies and landing zone operating models.
  • Demonstrated experience with incident, problem, change, and release management practices across complex and/or highly regulated environments.
  • Knowledge of cyber recovery concepts, including ransomware recovery, clean-room recovery, isolated recovery environments, backup integrity validation and cyber incident recovery planning.
  • Strong written and verbal communication skills, including demonstrating proficiency in producing clear executive reporting, operational dashboards and audit-ready evidence.
  • Bachelor’s degree in Computer Science, Engineering, Information Systems or equivalent.
  • P&C insurance domain experience, Guidewire/Snowflake exposure, AWS/Rubrik/Dynatrace certifications would be considered assets.

What you’ll get

  • The salary band for this position ranges from $125,000 – $175,000. Please note that individual salary is determined by factors such as job-related knowledge, skills and experience, as well as internal equity.
  • Compelling rewards package including base compensation, eligibility for annual bonus, retirement savings, share plan, health benefits, personal wellness, and volunteer opportunities.
  • Outstanding Career Development opportunities.
  • We’ll support your professional development education.
  • Competitive vacation package with the option to purchase 5 extra days off per year.
  • Employee driven programs focused on gender, LGBTQ+, origins, diversity, and inclusion.
  • Corporate wellness programs to support our employees’ physical and mental health.
  • Employee discount on home and auto insurance (where applicable).
  • Hybrid flexible work model.

Aviva Canada may use AI (Artificial Intelligence) tools to assist us throughout the recruitment process to screen, assess or select applicants for a position.

Aviva Canada welcomes applications from all qualified individuals and has a process in place to provide accommodations for persons with disabilities at all stages of the hiring process and during employment. If you require an accommodation during the interview or hiring process, please contact your Aviva Talent Acquisition Partner so that an appropriate accommodation can be arranged.

#LI-PS1

#LI-Hybrid

#J-18808-Ljbffr
Create a job alert for this search

Senior Manager, Site Reliability & Infrastructure Engineering • Markham, ON, Canada

Similar jobs

Manager, Site Reliability Engineering

MastercardToronto, Canada
Full-time

Lead Tubi's Site Reliability Engineering team as a Senior Manager, focusing on operational excellence and innovative practices.This hybrid role emphasizes resilience, automating systems, and dr... Show more

 • Promoted

Manager, Site Reliability Operations

Canadian Tire CorporationToronto
Full-time

Reporting to the AVP, Supply Chain Technology, SRE Operations, the Chapter Manager, SRE Development & Reliability, will be responsible for ensuring Supply Chain systems are operational and monitore... Show more

 • Promoted

Senior Site Project Manager – Transportation

R.V. Anderson Associates LimitedToronto, ON, CA
Full-time

A Canadian engineering firm is seeking a Senior Project Manager to join their Transportation team in Toronto.This full-time role requires a Civil Tech diploma or equivalent in Engineering, alongsid... Show more

 • Promoted

Senior Site Reliability Engineer

Guidewire SoftwareToronto, Ontario, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Senior Site Reliability Engineer - Hybrid & Leadership - C$120,000 - C$140,000 A Year

AgriTechEast York, Canada
Full-time

Senior Site Reliability Engineer needed to lead infrastructure, improve product resilience, and mentor a team in Vancouver.Requires 5+ years of cloud experience (AWS/GCP). Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Senior Transportation Design & Project Lead

Chisholm Fleming and AssociatesMarkham, York Region, CA
Full-time

A reputable engineering consultancy in York Region, Canada, is seeking a Mid-Senior level Civil or Transportation Engineer to join their team.This role involves managing transportation infrastructu... Show more

 • Promoted

Manager, Site Reliability Engineering - C$140,600 - C$190,600 A Year

Thomson ReutersToronto, Canada
Full-time

Our Privacy Statement & Cookie Policy**Manager, Site Reliability Engineering (SRE) page is loaded## Manager, Site Reliability Engineering (SRE)remote type:Hybridlocations:Canada, Toronto, Ontar... Show more

 • Promoted

Project Manager, Demolition

Green Infrastructure Partnersmarkham, on, Canada
Full-time

Come help build the future, where our work makes the world work!.Green Infrastructure Partners Inc.Reporting to the Senior Vice President, you will play a key role in delivering complex infrastruct... Show more

 • Promoted

Senior Site Reliability Engineer, Kong Konnect

Kong Inc.Toronto
Full-time

Senior Site Reliability Engineer, Kong Konnect.Join to apply for the Senior Site Reliability Engineer, Kong Konnect role at Kong Inc.Are you ready to power the World's connections?.If you don’t thi... Show more

 • Promoted

Engineering Manager, Transmission Line Projects

Stantec Consulting International Ltd.Markham, York region, Canada
Full-time

Manage a growing team of transmission line engineers.Ensure project execution aligns with industry standards in a flexible hybrid environment.As the Engineering Manager for Transmission Line Projec... Show more

 • Promoted

Deputy Project Manager – Conveyance & Underground Infrastructure

AECOMMarkham, ON, CA
Full-time

Deputy Project Manager – Conveyance & Underground Infrastructure.Primary Location: CA - Markham, ON - 105 Commerce Vall.Compensation: CAD 140,000 - CAD 180,000 - yearly.GTA Conveyance team (Mississ... Show more

 • Promoted

Site Reliability Engineer

Future Secure AIToronto
Full-time

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Show more

 • Promoted

Senior Manager, Site Reliability Engineering

DawninfotekToronto, Canada
Full-time

Senior Manager, SiteReliabilityEngineering(SRE) Contract to hire for a. Show more

 • Promoted

Senior Site Reliability Engineer

Magnet Forensicstoronto, on, Canada
Full-time

Magnet Forensics is a global leader in the development of digital investigative software that acquires, analyzes, and shares evidence from computers, smartphones, tablets, and IoT-related devices.O... Show more

 • Promoted

Lead Project Manager for Major Infrastructure

Green Infrastructure Partners Inc.Markham, ON, CA
Full-time

Grow your career as a Lead Project Manager with GIP in Markham, ON, where you can oversee significant infrastructure projects.Ensure projects meet all safety, budget, and quality standards with a c... Show more

 • Promoted

Senior Manager, Site Reliability Engineering - C$164,600 - C$235,100 A Year

TubiNorth York, Canada
Full-time

Lead and grow a Site Reliability Engineering team, focusing on platform resilience, automation, and AI integration.Drive technical strategy, operational excellence, and cross-functional collaboration. Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, Canada
Full-time

About the Team Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment D... Show more

 • Promoted

Senior Engineering Manager, Site Reliability - C$243,000 - C$297,000 A Year

RelayNorth York, Canada
Full-time

Leads the Site Reliability Engineering (SRE) team, defining strategy, improving platform reliability, performance, and resilience, and influencing engineering and product decisions. Show more