Talent.com
IFS
Senior Lead Site Reliability EngineerIFS • Vancouver, British Columbia, Canada
Senior Lead Site Reliability Engineer

Senior Lead Site Reliability Engineer

IFS • Vancouver, British Columbia, Canada
30+ days ago
Job type
  • Full-time
  • Permanent
Job description

Role Overview

  • As a Senior Lead Site Reliability Engineer (SRE) specializing in Azure you will be a hands-on technical owner of our cloud infrastructure. You will architect build and operate the systems that underpin our Azure-based SaaS offerings owning reliability scalability and security from the infrastructure layer up. You will work in close partnership with R&D to embed operational excellence into the software delivery lifecycle and you take full ownership of every system within Cloud Operations purview. You bring deep Azure and DevOps expertise thrive in complex distributed environments and raise the technical bar through the quality of your engineering work.

Key Responsibilities

  • Design implement and continuously improve Azure-based infrastructure for high-availability mission-critical SaaS services owning the full lifecycle from architecture through to production operation.
  • Own operate and continuously improve CI/CD pipelines across Jenkins Azure DevOps and GitHub Actions including pipeline architecture build performance deployment reliability secrets handling and migration work as we evolve our toolchain. This is active ownership not support.
  • Configure and maintain Ansible playbooks for configuration management provisioning automation and drift remediation across the infrastructure estate.
  • Build and maintain Infrastructure as Code using Terraform and/or ARM/Bicep covering the full provisioning lifecycle from initial environment build through to day-two operations and ongoing change management.
  • Work directly and continuously with R&D engineering teams to embed reliability operability and deployment quality into the software development lifecycle including pipeline design reviews pre-production environment ownership release readiness and incident learnings fed back into build practices.
  • Own the observability and alerting stack across Azure Monitor Log Analytics Application Insights Prometheus Grafana and Pingdom including metric collection synthetic monitoring coverage alerting thresholds and dashboard design. Own the PagerDuty configuration end-to-end: escalation policies routing rules service integrations and on-call schedule management. Act as the technical escalation point for complex incidents and participate in the teams on-call rotation.
  • Design operate and optimize AKS clusters for production workloads including node pool configuration autoscaling network policy ingress architecture workload identity and persistent storage patterns. Own cluster health upgrade lifecycle and capacity planning end-to-end.
  • Instrument Kubernetes workloads with Prometheus exporters and build Grafana dashboards that give engineering teams genuine operational visibility into service health latency error rates and resource consumption.
  • Take full technical ownership of all systems within Cloud Operations scope infrastructure tooling pipelines observability and security controls. If it lives in our environment you own its reliability its documentation and its improvement roadmap.
  • Lead root cause analysis on production incidents; author post-mortems with actionable engineering remediation not just process changes.
  • Define instrument and own SLOs SLIs and error budgets for Azure-hosted SaaS services; use data to drive reliability investment decisions.
  • Engineer and enforce security controls across identity access secrets and certificate management in Azure including hands-on implementation not just policy definition. Contribute directly to the technical controls evidence collection and continuous compliance posture required to maintain SOC 2 Type II ISO 27001 and ISO 9001 certification across the Cloud Operations environment.
  • Evaluate emerging Azure services and features against real production requirements; build proof-of-concepts validate at scale and drive adoption where the engineering case is clear.
  • Produce and maintain architecture documentation runbooks and operational playbooks that are technically precise enough for an on-call engineer to execute under pressure and meet the documentation standards required under our ISO 9001 quality management obligations.

Qualifications :

Required

  • 7 years in SRE Cloud Operations or DevOps roles with at least 4 years of hands-on Microsoft Azure focus.
  • Deep expertise across Azure services including App Services AKS Azure SQL Storage Networking Security Centre and Monitor.
  • Hands-on experience building maintaining and improving CI/CD pipelines in Jenkins Azure DevOps and GitHub Actions including real ownership of pipeline failures performance and evolution not just consumption.
  • Working experience with Ansible for configuration management and infrastructure automation.
  • Production-grade Kubernetes/AKS experience cluster operations workload troubleshooting RBAC network policies Helm and upgrade management in a live SaaS environment.
  • Hands-on experience with Prometheus and Grafana in a production context metric instrumentation alerting rule design and dashboard development not just consumption.
  • Experience with Pingdom for synthetic monitoring and PagerDuty for incident alerting and on-call management including configuration of escalation policies alert routing and participation in a 24/7 on-call rotation.
  • Strong scripting and automation skills in PowerShell Python Bash or equivalent with a track record of using code to eliminate operational toil.
  • Proven production-grade experience with Infrastructure as Code using Terraform and/or ARM/Bicep.
  • Advanced troubleshooting ability across distributed systems network layers and application performance in Azure comfortable owning a complex outage end-to-end.
  • Demonstrated ability to work closely and effectively with software development teams contributing to SDLC processes pipeline standards and release quality as a technical peer not a service desk.
  • Strong working knowledge of security protocols certificate lifecycle management secrets management and compliance controls in Azure including practical experience supporting or maintaining SOC 2 Type II ISO 27001 or ISO 9001 audits in an infrastructure or cloud operations context.
  • Demonstrated experience leading incident response and driving post-mortem remediation to completion.

Preferred

  • Azure certifications (Azure Solutions Architect Azure DevOps Engineer Expert or equivalent).
  • Experience with hybrid or multi-cloud environments including AWS.
  • Familiarity with Azure cost management tooling and hands-on optimisation work.
  • Experience operating large-scale SaaS platforms with multi-tenant infrastructure.
  • Experience with Grafana alerting Grafana OnCall or similar on-call routing tooling.


Additional Information :

What Were Offering

  • Salary Range: $133k and $151k CAD
  • Permanent Full-time

Use of Artificial Intelligence in Recruitment
As part of our recruitment process we may use automated tools including artificial intelligence to help screen and assess applications based on jobrelated criteria such as skills experience and qualifications.
These tools do not make hiring decisions. All employment decisions are reviewed and made by members of our hiring team.

We embrace flexibility and hybrid work opportunities to support diverse needs and lifestyles while also valuing inclusive workplace experiences. By fostering a sense of community we drive innovation strengthen connections and nurture belonging. Our commitment ensures you can work in a way that suits you best while also engaging with colleagues to share ideas and build meaningful relationships.


Remote Work :

No


Employment Type :

Full-time


Experience: years
Vacancy: 1
Create a job alert for this search

Senior Lead Site Reliability Engineer • Vancouver, British Columbia, Canada

Similar jobs

Senior Site Reliability Engineer

CerebrasVancouver, Metro Vancouver Regional District, Canada
Full-time

We’re seeking a senior Site Reliability Engineer/DevOps who is passionate about building the best infrastructure and maintaining the health of the systems.Design and maintain scalable, secure, and ... Show more

 • Promoted

Senior Site Reliability Engineer - $120,000 - $140,000 A Year

Semios GroupWest End, Canada
Full-time

Seeking a Senior Site Reliability Engineer to ensure infrastructure scalability, reliability, and performance.Focus on automation, observability, and resilience for SaaS products.Hands-on role with... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificVancouver, Metro Vancouver Regional District, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Senior Site Reliability Engineer - $113,900 - $147,400 A Year

SamsungVancouver, Canada
Full-time

Position Summary We're looking for a Senior SRE to lead our Cloud Engineering SmartThings team in driving the reliability, scalability, and performance of mission‑critical systems processing te... Show more

 • Promoted

Site Reliability Engineer Ii - C$100,000 - C$139,500 A Year

Electronic ArtsWest End, Canada
Full-time

The Site Reliability Engineer will design, build, and optimize systems.They will develop features and improve the platform for hosting EA's games and create automation tools. Show more

 • Promoted

Staff Site Reliability Engineer - C$124,200 - C$166,700 A Year

Walt Disney Animation StudiosVancouver, Canada
Full-time

Seeking a Staff Site Reliability Engineer with expertise in Linux, software development, CI tools, Git, cloud hosting, and container computing to optimize service deployments and improve system ava... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalVancouver, Metro Vancouver Regional District, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer

Arista NetworksVancouver, Canada
Full-time

Company DescriptionArista Networks is an industry leader in data‐driven, client‐to‐cloud networking for large data center, campus and routing environments.What sets us apart is our relentless pursu... Show more

 • Promoted

Site Reliability Engineer - $86,105 - $128,150 A Year

SmartThingsWest End, Canada
Full-time

Seeking an SRE/DevOps Engineer to maintain cloud systems, automate infrastructure, optimize workflows, and ensure operational excellence.Collaborate with dev teams and mentor junior engineers. Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseVancouver, Metro Vancouver Regional District, Canada
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsVancouver, Metro Vancouver Regional District, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Staff Site Reliability Engineer, Fabric

MongoDBVancouver, Canada
Full-time

The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Amon... Show more

 • Promoted

Site Reliability Engineer Vancouver, BC

LayerZerovancouver, metro vancouver regional district, Canada
Full-time

Founded in 2021, LayerZero’s vision is to create a community of cross-chain developers, building dApps that are no longer constrained by individual blockchain capabilities.With LayerZero's simple, ... Show more

 • Promoted

Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc.Vancouver, Metro Vancouver Regional District, CA
Full-time

A leading supply chain solutions provider is seeking a Site Reliability Engineer to optimize and ensure the reliability of their cloud infrastructure across AWS and Kubernetes.This role emphasizes ... Show more

 • Promoted

Site Reliability Engineer

AppleVancouver, Metro Vancouver Regional District, Canada
Full-time

The Apple Service Engineering - SRE team is looking for Site Reliability Engineers with experience in developing processes, tools, and automation for managing distributed systems in production envi... Show more

 • Promoted

Senior Site Reliability Engineer - Hybrid & Leadership - $120,000 - $140,000 A Year

A Leading Agricultural Technology CompanyWest End, Canada
Full-time

Seeking a Senior Site Reliability Engineer to lead infrastructure, improve product resilience, and mentor a team in Vancouver.Requires 5+ years of experience in cloud environments. Show more

 • Promoted

Senior Sre: Hybrid Cloud & Reliability Leader - C$120,000 - C$140,000 A Year

Leading Agricultural Technology CompanyVancouver, Canada
Full-time

Seeking a Senior Site Reliability Engineer to ensure service scalability and reliability in a hybrid cloud environment for an agricultural technology company.Requires expertise in cloud, automation... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsVancouver, Metro Vancouver Regional District, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Remote Site Reliability Engineer - Scale Crypto Systems

NewtonVancouver, Metro Vancouver Regional District, Canada
Remote
Full-time

A leading innovative tech company in Toronto is looking for a Site Reliability Engineer.In this pivotal role, you will enhance the reliability and resilience of critical services, manage incidents,... Show more

 • Promoted

Site Reliability Engineer

ArbitrumVancouver, British Columbia, Canada
Full-time

LayerZero The Future is Omnichain.Founded in 2021, LayerZero’s vision is to create a community of cross-chain developers, building dApps that are no longer constrained by individual blockchain capa... Show more