Talent.com
Socket.dev
Senior Observability & SRE LeaderSocket.dev • Toronto, ON, Canada
Senior Observability & SRE Leader

Senior Observability & SRE Leader

Socket.dev • Toronto, ON, Canada
3 days ago
Job type
  • Full-time
Job description

Marsh is seeking a visionary, transformational leader to reimagine and rebuild our Observability and Site Reliability Engineering function from the ground up. This is not a role for someone who wants to maintain the status quo. We need a leader who will fundamentally shift this function to a predictive, data-driven engineering discipline that prevents outages before they happen, embeds reliability into every system from design through production, and treats observability data as a strategic asset - not just an operational tool.
This is a career-defining opportunity to build a world-class observability and SRE organization at Fortune 500 scale.

Job Responsibilities

STRATEGIC VISION & PLATFORM TRANSFORMATION

  • Define and execute an observability and SRE strategy that shifts the organization from reactive operations to predictive reliability engineering.
  • Architect and deliver a unified, full-stack observability platform covering metrics, traces, logs, real-user monitoring (RUM), synthetic monitoring, and business-level KPIs - across on-prem, multi-cloud (AWS/Azure), containers, and SaaS integrations.
  • Rationalize and consolidate the current fragmented tooling landscape into a cohesive, cost-optimized platform. Eliminate redundant tools, reduce alert noise by 80%+, and establish a single pane of glass for system health.
  • Drive adoption of OpenTelemetry as the standard instrumentation framework, ensuring vendor-agnostic telemetry collection and future portability.

PREDICTIVE & PROACTIVE RELIABILITY

  • Build and operationalize AIOps and ML-driven capabilities to detect anomalies, predict failures, and surface emerging risks before they impact customers. Move beyond threshold-based alerting to intelligent, context-aware detection.
  • Establish automated correlation engines that link infrastructure signals, application traces, deployment events, and change records to dramatically reduce diagnostic time and identify root cause automatically.
  • Design and implement self-healing automation that detects, diagnoses, and remediates common failure patterns without human intervention – targeting 40%+ of recurring incidents for autonomous resolution.
  • Introduce chaos engineering and reliability testing programs (GameDays, fault injection, load testing) to proactively discover weaknesses before production incidents reveal them.

SITE RELIABILITY ENGINEERING CULTURE

  • Transform the existing operations-centric team into a modern SRE organization with embedded reliability engineers across product and platform squads, operating under a "you build it, you own it" model.
  • Define and implement SLO/SLI/Error Budget frameworks across critical services, creating a shared language between engineering, product, and business stakeholders for reliability decisions.
  • Drive the adoption of DevOps practices, CI/CD pipelines, and infrastructure as code using tools like Terraform or CloudFormation to manage infrastructure.
  • Champion reliability-first design principles - ensuring observability, graceful degradation, circuit breaking, and failure isolation are architected into every system from day one, not bolted on after launch.

INCIDENT PREVENTION & RAPID RECOVERY

  • Partner with Major Incident Management and Problem Management to build closed-loop feedback systems - every incident produces a reliability improvement, not just a postmortem document.
  • Drive MTTR toward minutes (not hours) through automated diagnostics, pre-built remediation playbooks, and intelligent correlation that tells responders what is wrong, not just that something is wrong.
  • Establish "Incidents Prevented" as a primary success metric alongside traditional MTTR/MTTD measures.

BUSINESS-ALIGNED OBSERVABILITY

  • Elevate observability from infrastructure metrics to business outcomes. Build real-time dashboards that connect system health to revenue impact, customer experience scores, and SLA compliance.
  • Integrate observability insights into ITSM (ServiceNow), data platforms, and executive reporting - making reliability data a first-class input to business and technology decision-making.

ENGINEERING & OPERATIONAL EXCELLENCE

  • Own the total cost of ownership of the observability platform. Optimize spend through data tiering, intelligent sampling, retention policies, and vendor negotiations. Deliver more insight per dollar.
  • Manage strategic vendor relationships (Datadog, Splunk, Logic Monitor, cloud-native tooling) with a focus on maximizing value extraction, not just license management.
  • Build a platform engineering mindset: observability capabilities are delivered as self-service products to engineering teams – instrumentation libraries, dashboard templates, alerting-as-code, SLO toolkits.

TEAM BUILDING & LEADERSHIP

  • Recruit, develop, and retain a world-class team of SRE engineers, observability platform engineers, data and performance engineers, and reliability analysts.
  • Establish an Observability & SRE Centre of Excellence that drives standards, best practices, and enablement across the global enterprise.
  • Foster a learning culture through internal tech talks, blameless postmortems, chaos engineering programs, and industry engagement.

Required Experience & Expertise

  • 15+ years in technology with 8+ years in progressively senior observability, SRE, or platform reliability leadership roles.
  • Demonstrated track record of transforming reactive monitoring organizations into proactive, engineering-driven SRE functions at enterprise scale (10,000+ employees, 1,000+ applications).
  • Deep expertise across the full observability stack: metrics (Prometheus, Datadog, CloudWatch), distributed tracing (Jaeger, OpenTelemetry, Datadog APM), log aggregation (Splunk, ELK, Datadog Logs), synthetic monitoring, and RUM.
  • Hands-on experience defining and operationalizing SLO/SLI/Error Budget frameworks that drive engineering prioritization and business alignment.
  • Proven experience building AIOps / ML-driven anomaly detection and automated remediation capabilities - not just evaluating vendor demos, but delivering production systems that prevent real incidents.
  • Strong background in chaos engineering, resilience testing, and reliability-by-design practices (circuit breakers, bulkheads, graceful degradation, retry/backoff patterns).
  • Experience operating across hybrid infrastructure: on-premises data centers, AWS, Azure, containerized workloads (Kubernetes), and SaaS platforms.
  • Demonstrated ability to drive cultural and organizational transformation across large, complex enterprises with multiple business units and hundreds of engineering squads.
  • Experience managing $5M+ observability platform budgets and optimizing total cost of ownership while expanding coverage and capability.
  • Executive communication skills - ability to present reliability strategy, risk posture, and investment cases to C-suite and board-level audiences.
  • Visionary thinker who can articulate a compelling future state and build the roadmap to get there - then execute relentlessly.

Marsh (NYSE: MRSH) is a global leader in risk, reinsurance and capital, people and investments, and management consulting, advising clients in 130 countries. With annual revenue of over $27 billion and more than 95,000 colleagues, Marsh helps build the confidence to thrive through the power of perspective. For more information, visit corporate.marsh.com, or follow us on LinkedIn and X.

Marsh is committed to embracing a diverse, inclusive and flexible work environment. We aim to attract and retain the best people and embrace diversity of age background, disability, ethnic origin, family duties, gender orientation or expression, marital status, nationality, parental status, personal or social status, political affiliation, race, religion and beliefs, sex/gender, sexual orientation or expression, skin color, or any other characteristic protected by applicable law. In accordance with the Accessibility for Ontarians with Disabilities Act, 2005, Marsh will provide a reasonable accommodation to employees and prospective employees to the point of undue hardship upon request and as required in respect of the individual’s particular restrictions and limitations. If you require a specific accommodation because of a disability or medical need, please contact reasonableaccommodations@marsh.com.

Marsh is committed to hybrid work, which includes the flexibility of working remotely and the collaboration, connections and professional development benefits of working together in the office. All Marsh colleagues are expected to be in their local office or working onsite with clients at least three days per week. Office-based teams will identify at least one “anchor day” per week on which their full team will be together in person.

This is a New position.

R_353128

#J-18808-Ljbffr
Create a job alert for this search

Senior Observability & SRE Leader • Toronto, ON, Canada

Similar jobs

Senior SRE Leader: Scale Reliability & Observability

RootlyToronto, ON, CA
Full-time

A fast-growing tech startup in Toronto is seeking an experienced Site Reliability Engineer.The role involves enhancing service performance, owning CI/CD pipelines, and building automation tools.Ide... Show more

 • Promoted

Senior SAP Ariba Sourcing Lead — Strategy & Implementation

TechDigital GroupToronto, ON, CA
Full-time

A leading technology firm in Canada is seeking a Lead SAP Ariba Sourcing professional to implement sourcing strategies and manage project objectives.This role involves configuring the Ariba Sourcin... Show more

 • Promoted

Lead Resilience R&D Partnerships, Canada

Initial Therapeutics, Inc.Toronto, ON, CA
Temporary

This role will lead Moderna’s R&D collaboration strategy in Canada, shaping and expanding the company’s footprint within the national life sciences ecosystem.The lead will define and execute a clea... Show more

 • Promoted

Senior SRE: Automation, Observability & Batch Performance

KyndrylToronto, ON, CA
Full-time

A global technology services provider is seeking a Site Reliability Engineer in Toronto to enhance the reliability and efficiency of critical batch workloads.This mid-senior level contract role emp... Show more

 • Promoted

SRE Observability Engineer

Tata Consultancy ServicesToronto
Full-time

Tata Consultancy Services (TCS) is an equal opportunity employer, and embraces diversity in race, nationality, ethnicity, gender, age, physical ability, neurodiversity, and sexual orientation, to c... Show more

 • Promoted

Senior Strategy and Operations Lead

Thomson ReutersToronto, ON, CA
Full-time

Senior Strategy and Operations Lead.Chief of Staff and partnering with senior marketing leaders to drive strategic planning, performance management, and operational excellence.This position owns ke... Show more

 • Promoted

Lead Observability Engineering - C$200,000 - C$235,000 A Year

RobinhoodToronto County, Canada
Full-time

A finance tech company seeks an Observability Engineering Lead in Toronto to manage a team, enhance system reliability, and integrate observability tools. Show more

 • Promoted

Senior Sre Leader: Ai-Powered Reliability & Observability - C$164,600 - C$235,100 A Year

Tubi, Inc.East York, Canada
Full-time

Senior SRE Leader sought to manage a team and ensure service availability and performance. Show more

 • Promoted

AI-Enabled Risk & Resilience Transformation Senior Manager

PwC Canadatoronto, on, Canada
Full-time

Join a dynamic team of talented risk professionals who are at the forefront of transforming how organizations navigate today’s most critical challenges across fast-paced risk landscapes.In our Fina... Show more

 • Promoted

Director, Process Excellence and Enablement

Xplore Inc.markham, on, Canada
Full-time

Canada’s fibre, 5G and satellite broadband company for rural living.Xplore is committed to the relentless pursuit of an improved broadband experience for all Canadians.Xplore is building a world‑cl... Show more

 • Promoted

Chapter Manager, SRE Development & Reliability

Canadian Tire CorporationToronto, ON, CA
Full-time

Reporting to the AVP, Supply Chain Technology, SRE Operations, the Chapter Manager, SRE Operations & Support, will be responsible for ensuring Supply Chain systems are operational and monitored.Thi... Show more

 • Promoted

Senior Strategy and Operations Lead

PVH (Tommy Hilfiger/Calvin Klein)Toronto, ON, CA
Full-time

Senior Strategy and Operations Lead.The Senior Strategy and Operations Lead plays a critical role in supporting the Chief of Staff and partnering with senior marketing leaders to drive strategic pl... Show more

 • Promoted

Senior SRE Focused on Automation and Cloud

Morningstar Credit Ratings, LLCToronto, ON, CA
Full-time

Explore a Senior Site Reliability Engineer role with Morningstar in Toronto, ON, emphasizing AWS and CI/CD automation.This hybrid position is key to ensuring robust investment data operations.You w... Show more

 • Promoted

Strategic Leadership Opportunity at Cossette

CossetteToronto, ON, CA
Full-time

Transform client relationships at Cossette as a Senior Strategist.Drive creative strategies and innovative solutions in a hybrid work model.As a Senior Strategist, you will develop and implement im... Show more

 • Promoted

Senior Manager, Resilience Engineering (Scenario Testing)

ScotiabankToronto
Full-time

Join a purpose driven winning team, committed to results, in an inclusive and high‑performing culture.As the Senior Manager, Resilience Engineering (Scenario Testing), you contribute to the global ... Show more

 • Promoted

Lead Observability Engineering - C$200,000 - C$235,000 A Year

Finance Technology CompanyEast York, Canada
Full-time

Seeking an engineering leader to manage a team building observability systems for a finance technology company.Responsibilities include design improvements, integration, and reliability enhancement. Show more

 • Promoted

Senior Nuclear Leader

Arcadismarkham, on, Canada
Full-time

Arcadis is seeking a Senior Nuclear Practice Leader to join our Radiation Risk Management Practice.You will provide technical leadership and mentorship, support the growth of the environmental radi... Show more

 • Promoted

Lead Electrification Solutions for Innovative Energy Projects

Cpus Engineering Staffing Solutions Inc.Toronto, ON, CA
Full-time

Drive the electrification landscape as the Head of Electrification Solutions.Oversee the development of cutting-edge energy projects, focusing on EV charging, solar integration, and battery storage... Show more

 • Promoted

Senior Business Continuity & Resilience Manager

Compunnel Inc.Toronto, ON, CA
Full-time

Business Continuity Manager with a strong background in Disaster Recovery and Project Management.This mid-senior level role requires 3-5+ years of experience and is critical in ensuring the resilie... Show more

 • Promoted

Intermediate/Senior Transmission Planning Power Systems Studies Engineer

Stantec Consulting International Ltd.markham, york region, Canada
Full-time

At Stantec, we’re leading the energy transition by combining the strength of a large company with the agility of specialized teams to think big and tackle the hardest challenges.Join us, and you’ll... Show more