Talent.com
McKesson
Sr. Observability EngineeringMcKesson • Mississauga, Peel Region, Canada
Sr. Observability Engineering

Sr. Observability Engineering

McKesson • Mississauga, Peel Region, Canada
15 days ago
Salary
CA$99,100.00 yearly
Job type
  • Full-time
Job description

About the role

McKesson is seeking a Senior Observability Engineer to join our Platform Engineering team. You will serve as a key engineer and practitioner of McKesson’s enterprise observability strategy, supporting the design, deployment, and continuous improvement of our monitoring and telemetry platforms across cloud and on‑premises infrastructure. You will operate with a high degree of autonomy, applying SRE principles and an engineering‑first mindset to ensure the reliability, performance, and availability of mission‑critical healthcare technology systems.

Observability Platform Ownership

  • Engineer, implement, and operate enterprise‑grade observability platforms including Dynatrace, LogicMonitor, Grafana, and Prometheus across multi‑cloud and hybrid environments.
  • Lead the adoption and integration of OpenTelemetry (OTel) standards for distributed tracing, metrics collection, and log correlation across engineering teams.
  • Manage and optimize Cisco ThousandEyes for network path visibility, internet performance monitoring, and end‑user experience insights.
  • Define and enforce observability‑as‑code practices, managing configurations through version‑controlled pipelines (e.g., Terraform, Helm, GitOps).
  • Evaluate, recommend, and pilot emerging observability technologies and methodologies to continuously advance McKesson’s monitoring capabilities.

SRE & Engineering Excellence

  • Apply Site Reliability Engineering (SRE) practices including error budgets, toil reduction, and chaos engineering to enhance system resilience.
  • Define, implement, and report on Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Key Performance Indicators (KPIs) in partnership with engineering and business stakeholders.
  • Develop and maintain alerting frameworks that minimize noise and maximize signal fidelity, reducing mean time to detect (MTTD) and mean time to resolve (MTTR).
  • Build self‑service observability tooling and dashboards that enable engineering teams to own their own reliability metrics.
  • Lead blameless post‑incident reviews (PIRs) and drive remediation actions to prevent recurrence.

Cloud & Infrastructure Monitoring

  • Design comprehensive observability strategies for cloud‑native workloads across AWS, Azure, and/or GCP, including containers (Kubernetes/EKS/AKS), serverless, and microservices architectures.
  • Instrument infrastructure, application, and business‑layer telemetry (logs, metrics, traces) to provide end‑to‑end visibility.
  • Manage network performance monitoring and synthetic testing via ThousandEyes to proactively identify and resolve connectivity and latency issues impacting end users.
  • Partner with DevOps, NetOps, and Security teams to integrate observability signals into CI/CD pipelines, ITSM workflows, and SOC operations.

AI & LLM Observability (Emerging Capability)

  • Extend the organization’s existing observability platforms to instrument AI‑powered applications and autonomous agent workflows, capturing the telemetry unique to non‑deterministic systems.
  • Implement OpenTelemetry GenAI semantic conventions to standardize how AI agent telemetry (traces, spans, token usage, tool invocations) is collected across frameworks and services.
  • Design and maintain tracing pipelines that capture the full agent execution chain from user intent and planner decisions through tool calls, retrieval steps, and model responses, enabling root‑cause analysis beyond traditional request/response tracing.
  • Define AI‑specific SLOs and cost budgets covering token consumption, hallucination rate, tool invocation success rate, and agent task completion rate.
  • Build dashboards and alerting for AI behavioral signals: model drift, performance degradation, prompt/context version changes, and guardrail violations.
  • Collaborate with data science, MLOps, and application teams to integrate observability signals into AI evaluation pipelines and continuous improvement feedback loops.
  • Implement security‑focused AI monitoring including prompt injection detection, PII exposure in logs, and anomalous tool call patterns in alignment with OWASP Top 10 for LLMs.

Collaboration & Technical Leadership

  • Act as an observability subject‑matter expert and trusted resource for engineers, architects, and product teams across McKesson.
  • Develop and maintain internal observability standards, runbooks, and best‑practice documentation.
  • Mentor and guide junior and mid‑level engineers on observability tooling, instrumentation techniques, and SRE principles.
  • Anticipate organizational scaling needs and proactively direct engineering efforts to stay ahead of capacity, reliability, and visibility challenges.
  • Contribute to the development of new frameworks, methodologies, and tooling that advance McKesson’s engineering maturity.

Required Qualifications

  • Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent practical experience.
  • 7+ years of experience in observability, monitoring, SRE, or platform engineering roles within complex, enterprise‑scale environments.
  • Hands‑on expertise with two or more of the following platforms: Dynatrace, LogicMonitor, Grafana, ThousandEyes, and Prometheus.
  • Demonstrated experience implementing OpenTelemetry (OTel) for distributed tracing and metrics instrumentation in polyglot service environments.
  • Proficiency with Cisco ThousandEyes or similar platforms for network monitoring, synthetic testing, and internet intelligence.
  • Deep understanding of cloud infrastructure monitoring across AWS, Azure, and/or GCP, including IaaS, PaaS, and container platforms.
  • Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management.
  • Scripting/automation proficiency in Python, Go, Bash, or equivalent, with experience building or extending monitoring integrations.
  • Experience with infrastructure‑as‑code and GitOps practices (Terraform, Helm, Ansible, etc.) as applied to observability configuration management.
  • Demonstrated ability to work independently, set direction, and deliver results in a complex, matrixed organization.

Preferred Qualifications

  • Experience in healthcare IT, regulated industries, or large‑scale enterprise environments.
  • Familiarity with AIOps platforms and ML‑based anomaly detection capabilities within Dynatrace or similar tooling.
  • Knowledge of eBPF‑based observability approaches and agent‑less instrumentation techniques.
  • Experience integrating observability data into ITSM platforms (ServiceNow, PagerDuty) and security workflows (SIEM).

Relocation

Relocation is not budgeted for this role.

Office Requirement

We are Flex and Connect with 2 days a week in office.

Compensation

$99,100 – $132,100

Equal Opportunity Employer

McKesson provides equal employment opportunities to applicants and employees, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, age, genetic information, or any other legally protected category. McKesson is committed to being an Equal Employment Opportunity Employer and offers opportunities to all job seekers, including job seekers with disabilities. If you need a reasonable accommodation to assist with your job search or application for employment, please contact us by sending an email to Disability_Accommodation@McKesson.com or Accessibility@mckesson.ca. Resumes or CVs submitted to this email box will not be accepted.

#J-18808-Ljbffr
Create a job alert for this search

Sr. Observability Engineering • Mississauga, Peel Region, Canada

Similar jobs

CANDU Reactivity Control Engineering Specialist

AtkinsRéalisMississauga, Peel Region, CA
Full-time

Unlock your potential as a CANDU Reactivity Control Engineering Specialist with AtkinsRéalis, applying your expertise to drive innovation in nuclear projects within a flexible work environment.In t... Show more

 • Promoted

IVVQ Test Specialist: Shipyard & Field Tests

ThalesMississauga, Peel Region, CA
Full-time

A leading defense and security company is seeking an IVVQ Test Specialist based in North Vancouver.This role involves planning and executing test events, writing test plans, and analyzing requireme... Show more

 • Promoted

Sr. Manager, Development Engineering

SmartCentres REITVaughan, York Region, CA
Full-time

Development Engineering, Senior Manager.Reporting to the Senior Director, Engineering, you will assist in engineering design, budgeting, development approvals, permitting and tendering for commerci... Show more

 • Promoted

Sr. Staff Analog Design Engineer

SemtechBurlington, Halton Region, CA
Full-time

Semtech’s Signal Integrity Products group designs analog and mixed-signal Integrated Circuits used in high performance optical and electrical networks.Our high-speed chips enable systems such as: h... Show more

 • Promoted

Advanced Metering Architect

CapgeminiMississauga, Peel Region, CA
Full-time

Enterprise Architect – Advanced Metering Infrastructure (AMI).We are seeking an experienced Enterprise Architect specializing in Advanced Metering Infrastructure (AMI) to lead our utilities transfo... Show more

 • Promoted

VP Operations – Precision Optics & Mechatronics Leader

Donnell ConsultingMississauga, Peel Region, CA
Full-time

A leading consulting firm in Canada is seeking a Vice President of Operations to drive innovative strategies and operational excellence.You will oversee end-to-end production processes and collabor... Show more

 • Promoted

Aerospace QA & Process Excellence Manager

Great Connections Employment ServicesMississauga, Peel Region, CA
Full-time

A leading aerospace company is looking for a QA Manager based in Mississauga to ensure compliance with its Quality Management System and promote quality across the organization.The ideal candidate ... Show more

 • Promoted

Analog Design, Sr Staff Engineer

SynopsysMississauga, Peel Region, CA
Full-time

In this role, you will work on the design, development, and refinement of Multi‑Gbps NRZ & PAM4 SERDES IP.You will be part of a fast‑growing analog and mixed‑signal R&D team developing high‑speed (... Show more

 • Promoted

Senior Engineering Manager

SynitiMississauga, Peel Region, CA
Full-time

Get AI-powered advice on this job and more exclusive features.Capgemini, tackles the hardest work in data for the world’s largest organizations.We combine intelligent software with deep data expert... Show more

 • Promoted

SH/FT Director of Development Engineering

Shift ParadigmMississauga, Peel Region, CA
Full-time

Join SH/FT as the Director of Development Engineering, where you will oversee impactful engineering initiatives remotely for top-tier brands.Drive innovation in software development with a focus on... Show more

 • Promoted

Sr. Observability Engineering

McKessonMississauga
Full-time

McKesson is seeking a Senior Observability Engineer to join our Platform Engineering team.You will serve as a key engineer and practitioner of McKesson’s enterprise observability strategy, supporti... Show more

 • Promoted

Strategic IoT Software Implementation Manager for Hybrid Engagements

CHEPMississauga, Peel Region, CA
Full-time

Shape customer success as an IoT Software Implementation Manager.Spearhead seamless digital solution deployments while enabling exceptional customer experiences in a hybrid work environment.This pi... Show more

 • Promoted

Director of Engineering — Platform & Reliability (Remote)

CliniaMississauga, Peel Region, CA
Remote
Full-time

A tech-driven health company in Canada is seeking a Director of Engineering to lead an engineering team of 25.You will manage delivery, ensure platform reliability, and set engineering standards wh... Show more

 • Promoted

Senior Director, Global Quality Eng - Semiconductors/Optics

Edison Smart®Burlington, Halton Region, CA
Full-time

A leading technology firm is seeking a Senior Director of Quality Engineering in Burlington, Canada.The role involves developing and executing quality strategies in the semiconductor field, necessi... Show more

 • Promoted

RQ10361 - Sr. Project Manager/Leader (Technology Programs, Azure, Migrations, On prem to Cloud)

Source CodeMississauga, Peel Region, CA
Full-time

Project Manager/Leader (Technology Programs, Azure, Migrations, On prem to Cloud).To respond to the priorities of the Government of Ontario through the passage of the Bill 5, Protect Ontario by Unl... Show more

 • Promoted

Senior Radiation Risk Management Leader

Arcadisoakville, on, Canada
Full-time

Elevate your expertise at Arcadis by becoming a Senior Nuclear Practice Leader focusing on radiation risk and compliance.This role emphasizes leadership and mentorship within a growing team.As a Se... Show more

 • Promoted

Supplier Quality Engineering Specialist, Manufacturing

Siemens MobilityVaughan, York Region, CA
Full-time

Supplier Quality Engineering Specialist, Manufacturing.Contributes to the definition of the cross-functional commodity strategy and implements the Supplier Quality management (SQM) strategy and SQM... Show more

 • Promoted

Sr. QA Manual Testing - Canada (Remote)

RELQ TECHNOLOGIESMississauga, Peel Region, CA
Remote
Full-time

Client Domain – Product-based company - HRMS (Human Resource Management System) applications focused on SaaS-based HRMS solutions for the public sector.The ideal candidate has experience working in... Show more

 • Promoted

Sr Advanced Project Engineer – New Product Development

Honeywell AerospaceMississauga, Peel Region, CA
Full-time

As a Sr Advanced Project Engineer at Honeywell, you will lead and manage advanced engineering projects within the Aerospace business unit.You will provide technical expertise, project management, a... Show more

 • Promoted

Sr Manager, Platform & Integration Support

Ingram MicroMississauga, Peel Region, CA
Full-time

Senior Manager – AI‑Enabled Platform Support.Ingram Micro Xvantage™ ecosystem: Xvantage for Vendors (X4V) and Xvantage Integrations (XI).The position focuses on transforming support from reactive t... Show more