Talent.com
OMERS
Lead, Site Reliability Engineering (Application Support)OMERS • Toronto, ON, Canada
Lead, Site Reliability Engineering (Application Support)

Lead, Site Reliability Engineering (Application Support)

OMERS • Toronto, ON, Canada
13 days ago
Salary
CA$86,000.00 yearly
Job type
  • Full-time
Job description

Choose a workplace that empowers your impact.

Join a global workplace where employees thrive. One that embraces diversity of thought, expertise and experience. A place where you can personalize your employee journey to be - and deliver - your best.

We are a purpose-driven, dynamic and sustainable pension plan. An industry leading global investor with teams in Toronto to London, New York, Singapore, Sydney and other major cities across North America and Europe. We embody the values of our 665,000 members, placing their best interests at the heart of everything we do.

Join us to accelerate your growth & development, prioritize wellness, build connections, and support the communities where we live and work.

Don’t just work anywhere - come build tomorrow together with us.

Role Summary

The Lead, Site Reliability Engineering ensures monitoring and analysis is conducted to guarantee the ongoing stability of all systems. The Lead is also expected to implement and maintain the infrastructure and tools that are necessary to manage the software development process.

Reporting to the SRE, Developer Platform Engineering, the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal applications. This position combines platform engineering, site reliability engineering, and production support responsibilities to ensure applications remain reliable, secure, and easy to operate.

This is an excellent opportunity for someone who is passionate about Site Reliability Engineering and enjoys improving the reliability, availability, performance, and operability of cloud-native platforms. The successful candidate will apply SRE practices such as incident response, observability, automation, runbook development, root-cause analysis, and continuous reliability improvement while partnering with cross-functional teams to onboard, deploy, monitor, and support applications across the organization.

You Will Be Responsible For

  • Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments.
  • Respond to incidents, lead triage activities, and work closely with SRE, Platform, Network, Security, and application teams to restore service and resolve issues.
  • Support deployments, release activities, change management, and CI/CD pipelines, including GitHub Actions workflows.
  • Check, troubleshoot, and resolve user and developer access issues, including Azure AD groups, SSO, application permissions, and firewall rules.
  • Configure and support platform components such as Azure Container Apps, App Registrations, Key Vault, DNS, certificates, networking, and shared cloud services.
  • Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, Log Analytics, and related observability tools.
  • Support onboarding of new applications and teams to the DEV platform by helping with setup, access, deployment readiness, monitoring, and operational handover.
  • Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation to improve support effectiveness and knowledge sharing.
  • Contribute to automation and continuous improvement initiatives that reduce manual effort, improve reliability, and strengthen operational processes.
  • Provide technical guidance to team members and stakeholders while promoting Site Reliability Engineering and platform support best practices.

Required Skills & Experience

  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support.
  • Strong hands‑on experience with Microsoft Azure services, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Storage Accounts, Azure SQL, API Management (APIM), and Azure Functions.
  • Experience supporting production environments, including incident response, troubleshooting, problem management, and operational support processes.
  • Experience with container technologies and cloud-native application architectures.
  • Hands‑on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms.
  • Strong understanding of identity, networking, and access management concepts, including SSO, OAuth, application registrations, and security groups.
  • Experience with observability and monitoring platforms such as Datadog, Azure Monitor, and Log Analytics.
  • Understanding of cloud networking concepts, including DNS, certificates, firewalls, private endpoints, and network security controls.
  • Experience with scripting and automation using technologies such as PowerShell, Bash, Azure CLI, Python, or similar tools.
  • Strong knowledge of operating systems and cloud infrastructure concepts.
  • Proven ability to work effectively in cross-functional environments and collaborate with technical and business stakeholders.
  • Strong communication, problem‑solving, and organizational skills.

Preferred Skills & Experience

  • Experience with container apps and Kubernetes container orchestration platforms.
  • Experience with different pipelines.
  • Experience supporting enterprise developer platforms or internal platform engineering teams.
  • Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and reliability engineering practices.
  • Experience with enterprise API integrations and platform services.
  • Exposure to AI, LLM, or agent-based technology platforms.
  • Experience with Azure networking and security best practices in enterprise environments.
  • Azure, Network, DevOps, or cloud-related certifications.
  • Post‑secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline.

We believe that time together in the office is important for OMERS and Oxford, the strength of our employees, and the work we do for our pension members. In delivering on our pension promise, keeping us connected to our work and each other, our flexible hybrid work guideline requires teams to come in to the office 4 days per week.

This posting is for an existing vacancy.

The expected salary range for this position is $86,000.00 - $130,000.00 per year.

You may also be eligible to receive an annual Incentive Award pursuant to our Short-term Incentive plan and our Long-Term Incentive plan (if applicable), and to participate in our group benefits and retirement plans – details on these elements of compensation are included within OMERS & Oxford offer letters.

As one of Canada’s largest defined benefit pension plans, our people-first culture is at its best when our workforce reflects the communities where we live and work — and the members we proudly serve.

From hire to retire, we are an equal opportunity employer committed to an inclusive, barrier‑free recruitment and selection process that extends all the way through your employee experience. This sense of belonging and connection is cultivated up, down and across our global organization thanks to our vast network of Employee Resource Groups with executive leader sponsorship, our Purpose@Work committee and employee recognition programs.

Artificial intelligence (AI) tools are used to support certain stages of the OMERS recruitment process. While AI assists us in our process, human judgment and decision-making remain central to our candidate experience.

#J-18808-Ljbffr
Create a job alert for this search

Lead, Site Reliability Engineering (Application Support) • Toronto, ON, Canada

Similar jobs

Manager, Site Reliability Engineering

MastercardToronto
Full-time

Mastercard powers economies and empowers people in 200+ countries and territories worldwide.Together with our customers, we’re helping build a sustainable economy where everyone can prosper.We supp... Show more

 • Promoted

Senior Practice Lead – Application Release Engineering - C$115,600 - C$163,200 A Year

TD SecuritiesEast York, Canada
Full-time

Leading and developing a team, managing application release processes, and driving continuous improvement and automation. Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Automation Support Analyst (Site Reliability Engineering)

OMERSToronto, ON, CA
Full-time

We are looking for a curious, solutions-oriented Automation Support Analyst to join the Data and Technology team.This role is ideal for someone who enjoys solving production issues, working with mo... Show more

 • Promoted

Remote Enterprise Systems & Integration Lead - C$110,000 - C$147,000 A Year - Remote

Global Technology Solutions ProviderToronto County, Canada
Remote
Full-time

Lead enterprise systems and integration architecture, ensuring reliability and scalability of IT platforms and integration strategies remotely. Show more

 • Promoted

Remote Director, Site Reliability Engineering - C$238,000 - C$298,000 A Year - Remote

AffirmToronto County, Canada
Remote
Full-time

Seeking a Director of Site Reliability Engineering to lead a team, ensure high service availability, and drive operational excellence in a financial technology company. Show more

 • Promoted

Lead Site Reliability Engineer - C$140,000 - C$178,000 A Year

EpamToronto County, Canada
Full-time

Lead Site Reliability Engineer to ensure system reliability, observability, and performance monitoring for digital trading products.Responsible for leading monitoring initiatives, defining reliabil... Show more

 • Promoted

Remote Enterprise Systems & Integration Lead - C$110,000 - C$147,000 A Year - Remote

Global technology solutions providerToronto County, Canada
Remote
Full-time

Lead enterprise systems and integration architecture, ensuring system reliability and scalability.Requires expertise in Dynamics 365 CE, Power Apps, and Azure technologies. Show more

 • Promoted

Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc.Toronto, ON, CA
Full-time

A leading supply chain solutions provider is seeking a Site Reliability Engineer to optimize and ensure the reliability of their cloud infrastructure across AWS and Kubernetes.This role emphasizes ... Show more

 • Promoted

Configuration Lead Role - C$56,000 - C$79,400 A Year

CdwEast York, Canada
Full-time

Lead configuration specialist responsible for device imaging, provisioning, and technical support, ensuring quality service and process improvement within a warehouse facility.Requires strong custo... Show more

 • Promoted

Technical Lead for Application Deployments

CprvisionStouffville, Ontario, Canada
Full-time

Drive deployment success at Portfolio+ as a Technical Deployment Lead.Apply your skills in CI/CD pipeline management and cloud environments to ensure seamless application releases.As the Technical ... Show more

 • Promoted

Director of Engineering — Platform & Reliability (Remote)

CliniaToronto, ON, CA
Remote
Full-time

A tech-driven health company in Canada is seeking a Director of Engineering to lead an engineering team of 25.You will manage delivery, ensure platform reliability, and set engineering standards wh... Show more

 • Promoted

IBM Site Reliability Engineering Expert

LeadingtalentMarkham
Full-time

Step into a career as a Site Reliability Engineer at IBM, focused on enhancing system reliability and performance.Engage directly with production systems and optimize customer experience.In this ro... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Applications Systems Lead - C$100,000 - C$110,000 A Year

Pacific National ExhibitionNorth York, Canada
Full-time

Lead the implementation and maintenance of the Momentus ERP system, architecting a data warehouse and central dashboard for strategic decision-making.Manage integrations, vendor relationships, and ... Show more

 • Promoted

Applications Engineering Specialist (On-Site) - $65,000 - $75,000 A Year

G&W ElectricToronto, Canada
Full-time

This role involves interpreting customer specifications and selecting appropriate products, developing custom solutions, and building customer relationships.Requires an engineering degree and 1-3 y... Show more

 • Promoted

Manager, Site Reliability Engineering - C$140,600 - C$190,600 A Year

Thomson ReutersToronto, Canada
Full-time

Lead a Site Reliability Engineering team, focusing on system reliability, performance, automation, and DevOps practices.Drive strategic vision, operational excellence, and risk management for cloud... Show more

 • Promoted

Application Tech Lead

Smart IT Frame LLCToronto
Full-time

Mount Laurel NJ office 2-3 times a week or Toronto office 4 times a week.At Smart IT Frame, we connect top talent with leading organizations across the USA.With over a decade of staffing excellence... Show more

 • Promoted

Enterprise Solutions Engineering Lead - Remote

VerkadaToronto, ON, CA
Remote
Full-time

Verkada invites applications for the Remote Enterprise Solutions Engineering Manager position.This role is essential for guiding our engineering team toward excellence in technical solutions.As the... Show more

 • Promoted

Reliability Engineering Lead (Contract) - $37.4 - $58.4 An Hour

CapgeminiEast York, Canada
Full-time

Lead equipment and building reliability strategies, optimize maintenance, and perform root cause analysis for breakdowns in the commercial aircraft program. Show more