Talent.com
IFS
Senior Lead Site Reliability Engineer | IFS CopperleafIFS • Vancouver, British Columbia, CA
Senior Lead Site Reliability Engineer | IFS Copperleaf

Senior Lead Site Reliability Engineer | IFS Copperleaf

IFS • Vancouver, British Columbia, CA
26 days ago
Job type
  • Full-time
Job description

Job Description

Role Overview

  • As a Senior Lead Site Reliability Engineer (SRE) specializing in Azure, you will be a hands-on technical owner of our cloud infrastructure. You will architect, build, and operate the systems that underpin our Azure-based SaaS offerings — owning reliability, scalability, and security from the infrastructure layer up. You will work in close partnership with R&D to embed operational excellence into the software delivery lifecycle, and you take full ownership of every system within Cloud Operations' purview. You bring deep Azure and DevOps expertise, thrive in complex distributed environments, and raise the technical bar through the quality of your engineering work.

Key Responsibilities

  • Design, implement, and continuously improve Azure-based infrastructure for high-availability, mission-critical SaaS services — owning the full lifecycle from architecture through to production operation.
  • Own, operate, and continuously improve CI/CD pipelines across Jenkins, Azure DevOps, and GitHub Actions — including pipeline architecture, build performance, deployment reliability, secrets handling, and migration work as we evolve our toolchain. This is active ownership, not support.
  • Configure and maintain Ansible playbooks for configuration management, provisioning automation, and drift remediation across the infrastructure estate.
  • Build and maintain Infrastructure as Code using Terraform and/or ARM/Bicep, covering the full provisioning lifecycle — from initial environment build through to day-two operations and ongoing change management.
  • Work directly and continuously with R&D engineering teams to embed reliability, operability, and deployment quality into the software development lifecycle — including pipeline design reviews, pre-production environment ownership, release readiness, and incident learnings fed back into build practices.
  • Own the observability and alerting stack across Azure Monitor, Log Analytics, Application Insights, Prometheus, Grafana, and Pingdom — including metric collection, synthetic monitoring coverage, alerting thresholds, and dashboard design. Own the PagerDuty configuration end-to-end: escalation policies, routing rules, service integrations, and on-call schedule management. Act as the technical escalation point for complex incidents and participate in the team's on-call rotation.
  • Design, operate, and optimize AKS clusters for production workloads — including node pool configuration, autoscaling, network policy, ingress architecture, workload identity, and persistent storage patterns. Own cluster health, upgrade lifecycle, and capacity planning end-to-end.
  • Instrument Kubernetes workloads with Prometheus exporters and build Grafana dashboards that give engineering teams genuine operational visibility into service health, latency, error rates, and resource consumption.
  • Take full technical ownership of all systems within Cloud Operations' scope — infrastructure, tooling, pipelines, observability, and security controls. If it lives in our environment, you own its reliability, its documentation, and its improvement roadmap.
  • Lead root cause analysis on production incidents; author post-mortems with actionable engineering remediation, not just process changes.
  • Define, instrument, and own SLOs, SLIs, and error budgets for Azure-hosted SaaS services; use data to drive reliability investment decisions.
  • Engineer and enforce security controls across identity, access, secrets, and certificate management in Azure — including hands-on implementation, not just policy definition. Contribute directly to the technical controls, evidence collection, and continuous compliance posture required to maintain SOC 2 Type II, ISO 27001, and ISO 9001 certification across the Cloud Operations environment.
  • Evaluate emerging Azure services and features against real production requirements; build proof-of-concepts, validate at scale, and drive adoption where the engineering case is clear.
  • Produce and maintain architecture documentation, runbooks, and operational playbooks that are technically precise enough for an on-call engineer to execute under pressure — and meet the documentation standards required under our ISO 9001 quality management obligations.

Qualifications

Required

  • 7+ years in SRE, Cloud Operations, or DevOps roles, with at least 4 years of hands-on Microsoft Azure focus.
  • Deep expertise across Azure services including App Services, AKS, Azure SQL, Storage, Networking, Security Centre, and Monitor.
  • Hands-on experience building, maintaining, and improving CI/CD pipelines in Jenkins, Azure DevOps, and GitHub Actions — including real ownership of pipeline failures, performance, and evolution, not just consumption.
  • Working experience with Ansible for configuration management and infrastructure automation.
  • Production-grade Kubernetes/AKS experience — cluster operations, workload troubleshooting, RBAC, network policies, Helm, and upgrade management in a live SaaS environment.
  • Hands-on experience with Prometheus and Grafana in a production context — metric instrumentation, alerting rule design, and dashboard development, not just consumption.
  • Experience with Pingdom for synthetic monitoring and PagerDuty for incident alerting and on-call management — including configuration of escalation policies, alert routing, and participation in a 24/7 on-call rotation.
  • Strong scripting and automation skills in PowerShell, Python, Bash, or equivalent — with a track record of using code to eliminate operational toil.
  • Proven, production-grade experience with Infrastructure as Code using Terraform and/or ARM/Bicep.
  • Advanced troubleshooting ability across distributed systems, network layers, and application performance in Azure — comfortable owning a complex outage end-to-end.
  • Demonstrated ability to work closely and effectively with software development teams — contributing to SDLC processes, pipeline standards, and release quality as a technical peer, not a service desk.
  • Strong working knowledge of security protocols, certificate lifecycle management, secrets management, and compliance controls in Azure — including practical experience supporting or maintaining SOC 2 Type II, ISO 27001, or ISO 9001 audits in an infrastructure or cloud operations context.
  • Demonstrated experience leading incident response and driving post-mortem remediation to completion.

Preferred

  • Azure certifications (Azure Solutions Architect, Azure DevOps Engineer Expert, or equivalent).
  • Experience with hybrid or multi-cloud environments, including AWS.
  • Familiarity with Azure cost management tooling and hands-on optimisation work.
  • Experience operating large-scale SaaS platforms with multi-tenant infrastructure.
  • Experience with Grafana alerting, Grafana OnCall, or similar on-call routing tooling.

Create a job alert for this search

Senior Lead Site Reliability Engineer | IFS Copperleaf • Vancouver, British Columbia, CA

Similar jobs

Senior Site Reliability Engineer

CerebrasVancouver, Metro Vancouver Regional District, Canada
Full-time

We’re seeking a senior Site Reliability Engineer/DevOps who is passionate about building the best infrastructure and maintaining the health of the systems.Design and maintain scalable, secure, and ... Show more

 • Promoted

Principal Site Reliability Engineer

SaviyntVancouver, Canada
Full-time

Why Join Saviynt Work on a mission‑critical SaaS platform used by global enterprises Solve complex reliability challenges at scale Influence architecture and engineering culture at a company lev... Show more

 • Promoted

Senior On-Site Construction Leader – Infra & Utilities

Cross Fraser PartnershipVancouver, Metro Vancouver Regional District, CA
Full-time

A leading infrastructure development firm in Metro Vancouver is seeking a Senior Construction Engineer to oversee on-site activities for a significant infrastructure project.This position requires ... Show more

 • Promoted

Site-Based Reliability & Integrity Engineer – Lng Facility - C$115,000 - C$125,000 A Year

Woodfibre Management LtdSquamish, Canada
Full-time

A Canadian LNG project company is seeking a Reliability & Integrity Engineer for their facility in Squamish, BC.This role involves ensuring technical integrity and reliability during constructi... Show more

 • Promoted

Site-Based Reliability & Integrity Engineer – Lng Facility - C$115,000 - C$125,000 A Year

Canadian LNG project companySquamish, Canada
Full-time

A Canadian LNG project company is seeking a Reliability & Integrity Engineer for their facility in Squamish, BC.This role involves ensuring technical integrity and reliability during constructi... Show more

 • Promoted

Lead Platform Engineer Enhancing DevOps and System Reliability

Lillio (formerly HiMama)Vancouver, Metro Vancouver Regional District, CA
Full-time

Transform early childhood education as a Senior Platform Engineer focused on system performance and collaborative tooling.Drive key initiatives for scalable, reliable digital platforms.In this pivo... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalVancouver, Metro Vancouver Regional District, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseVancouver, Metro Vancouver Regional District, Canada
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Director of Engineering — Platform & Reliability (Remote)

CliniaVancouver, Metro Vancouver Regional District, CA
Remote
Full-time

A tech-driven health company in Canada is seeking a Director of Engineering to lead an engineering team of 25.You will manage delivery, ensure platform reliability, and set engineering standards wh... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificVancouver, Metro Vancouver Regional District, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Modernization Site Lead — Construction & Safety

KONEDelta, Metro Vancouver Regional District, CA
Full-time

A global leader in building solutions is seeking an experienced construction site leader in Delta, Canada.The role involves managing a team and subcontractors, implementing safety protocols, and en... Show more

 • Promoted

Site Reliability Engineer Vancouver, BC

LayerZerovancouver, metro vancouver regional district, Canada
Full-time

Founded in 2021, LayerZero’s vision is to create a community of cross-chain developers, building dApps that are no longer constrained by individual blockchain capabilities.With LayerZero's simple, ... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsVancouver, Metro Vancouver Regional District, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Senior Contaminated Sites Lead - Remote-Eligible

Stantec Consulting International Ltd.Vancouver, Metro Vancouver Regional District, CA
Remote
Full-time

A global engineering firm seeks a Senior Environmental Professional to lead environmental investigations and remediation projects across Canada.The role involves project management, client relation... Show more

 • Promoted

Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc.Vancouver, Metro Vancouver Regional District, CA
Full-time

A leading supply chain solutions provider is seeking a Site Reliability Engineer to optimize and ensure the reliability of their cloud infrastructure across AWS and Kubernetes.This role emphasizes ... Show more

 • Promoted

Site Reliability Engineer

AppleVancouver, Metro Vancouver Regional District, Canada
Full-time

The Apple Service Engineering - SRE team is looking for Site Reliability Engineers with experience in developing processes, tools, and automation for managing distributed systems in production envi... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsVancouver, Metro Vancouver Regional District, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Infrastructure Reliability Engineer

ShippoVancouver, Metro Vancouver Regional District, CA
Full-time

Enhance shipping solutions as a Senior Site Reliability Engineer in a remote setting.Focus on infrastructure integrity, scalability, and performance in a collaborative environment.This position inv... Show more

 • Promoted

Remote Site Reliability Engineer - Scale Crypto Systems

NewtonVancouver, Metro Vancouver Regional District, Canada
Remote
Full-time

A leading innovative tech company in Toronto is looking for a Site Reliability Engineer.In this pivotal role, you will enhance the reliability and resilience of critical services, manage incidents,... Show more

 • Promoted

Site Reliability Engineer

ArbitrumVancouver, British Columbia, Canada
Full-time

LayerZero The Future is Omnichain.Founded in 2021, LayerZero’s vision is to create a community of cross-chain developers, building dApps that are no longer constrained by individual blockchain capa... Show more