Talent.com
Shore Consulting
Lead Platform Engineer (DevOps & MLOps)Shore Consulting • Toronto, Ontario, Canada
Lead Platform Engineer (DevOps & MLOps)

Lead Platform Engineer (DevOps & MLOps)

Shore Consulting • Toronto, Ontario, Canada
2 days ago
Job type
  • Full-time
Job description

Job Description

Reporting to the Director, Platform Services, the Lead DevOps Engineer is a senior, hands-on technical leader responsible for building, operating, and continuously improving Afflo’s production and non-production cloud environments, CI/CD pipelines, observability, and operational toolchain. This role translates the VP’s operational strategy, standards, and compliance objectives into reliable, scalable, secure implementation while leading execution across the DevOps function day to day.

You will work closely with Product Engineering, QA, Service Management, Implementation/Project Delivery, and external vendors to ensure Afflo services meet uptime, performance, security, and audit expectations in regulated healthcare contexts. You will also mentor other DevOps engineers, lead incident response and prevention work, and drive practical improvements that reduce operational risk and accelerate safe delivery.

This role is demanding and diverse, involving:

  • Operational ownership of cloud infrastructure and delivery pipelines

  • Release engineering and environment lifecycle management

  • Observability, incident leadership, and continuous improvement

  • Security controls, evidence readiness, and DR/BCP execution

  • Tooling automation that reduces toil and improves team productivity

Responsibilities

Operational Ownership

  • Own the reliability and day-to-day operation of Afflo environments (production and non-production), ensuring uptime, performance, responsiveness, and strong operational hygiene.

  • Lead triage, mitigation, and restoration during incidents; coordinate with Service Management and engineering stakeholders through resolution.

  • Conduct and author post-incident reviews and drive prevention work to reduce recurrence, improve MTTR, and increase change safety.

  • Establish and maintain on-call standards, escalation paths, maintenance practices, and operational runbooks aligned with IT Operations and System Administration policies.

Cloud Infrastructure Engineering (IaC-First)

  • Design, build, and maintain secure, resilient cloud infrastructure using Infrastructure as Code (IaC) with reusable modules, review discipline, and predictable environment patterns.

  • Build and improve environment lifecycle workflows (provision, reset, clone, teardown) for QA/UAT/demo/customer environments and internal team needs.

  • Implement secure-by-default patterns: network segmentation, least privilege, secrets handling, encryption, audit logging, and access reviews.

  • Perform capacity planning and cost optimization—balancing availability, scalability, and operating cost, and providing actionable recommendations to the VP of Delivery.

  • Design, set up and maintain AI specific workloads and pipelines. E.g. Data processing, model training, inference etc.

CI/CD, Release Engineering, and Delivery Enablement

  • Build and maintain automated CI/CD pipelines to enable rapid, safe deployments, including release gates, automated checks, artifact integrity, and rollback readiness.

  • Participate in and/or lead major release windows and maintenance deployments; ensure readiness checks, comms coordination, and post-release verification.

  • Standardize release processes across teams/products to reduce variance, improve predictability, and support project timelines and SLAs.

  • Partner with Product Engineering and QA to improve test reliability, deployment quality, and developer experience.

Observability and Monitoring

  • Implement and maintain monitoring, alerting, logging, and dashboards that provide actionable signals for availability, performance, security, and data integrity.

  • Reduce alert noise and improve detection coverage through tuning, SLO/SLI development, and automated verification checks.

  • Provide operational insights to engineering teams using logs and metrics to identify trends, performance constraints, and failure patterns.

Security, Compliance, and Audit Readiness Enablement

  • Implement operational controls and evidence-producing mechanisms aligned with IT Operations policies and selected frameworks (e.g., SOC2/ISO-aligned practices).

  • Support security and governance requests by producing operational materials (diagrams, environment descriptions, safeguards, maintenance practices) and operational evidence in a timely manner.

  • Coordinate with Security/Service Management on vulnerability management, patching practices, vendor security events, and operational monitoring requirements.

  • Contribute to disaster recovery and business continuity readiness by maintaining runbooks, validating backups/restores, and participating in recovery exercises/tabletop tests.

Tooling, Internal Enablement, and Cross-Team Support

  • Support onboarding/offboarding and access provisioning across enterprise tools (email, document storage, chat, ticketing, VPN, dev/QA environment access), emphasizing least privilege and traceability.

  • Build automation/scripts to streamline frequent employee tasks and reduce operational toil.

  • Maintain a clear, prioritized operational ticket pipeline; triage requests, track outcomes, and communicate progress and risks.

Vendor Collaboration and Operational Toolchain

  • Work with vendors and internal stakeholders to procure, configure, and maintain operational tooling (hosting, monitoring, backups, authentication services, pipeline tools).

  • Coordinate vendor-driven maintenance/outages and ensure internal and customer-facing communications occur when required.

  • Provide practical input on tool selection and implementation feasibility, aligned to the VP’s standards and roadmap.

Leadership Within the DevOps Function

  • Mentor DevOps engineers through pairing, code/IaC reviews, incident coaching, and documentation/runbook development.

  • Raise team maturity by defining “how we do it here”: templates, standards, checklists, guardrails, and repeatable operational processes.

  • Serve as senior escalation for complex infrastructure/pipeline issues and lead cross-team problem-solving efforts.


Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).

  • 5–10 years of progressive experience in DevOps/SRE/infrastructure engineering, including ownership of production systems.

  • Strong Linux, networking, and troubleshooting skills across distributed systems.

  • Advanced experience with cloud environments (Azure and/or GCP preferred; multi-cloud exposure is an asset).

  • Expert-level Infrastructure as Code experience (e.g., Terraform/Pulumi), including modular design, review practices, and safe change management.

  • Strong Kubernetes experience (operations, deployments, security posture, cluster/platform troubleshooting).

  • Strong experience with designing and building AI training and inference workflows in a cloud environment.

  • Proven CI/CD and release engineering experience (e.g., GitLab CI, Jenkins, ArgoCD or equivalent), including quality gates and safe deployment strategies.

  • Proven experience with software development lifecycle (SDLC) methodologies and best practices,

  • Experience with IT Service Management (ITSM) (ServiceNow, JIRA Service Management, BNC Remedy) and Kanban project management (JIRA Software or equivalent).

  • Demonstrated incident leadership (on-call participation, incident coordination, RCA authorship, prevention follow-through).

  • Security-minded approach: least privilege, secrets management, vulnerability management, audit logging, and regulated-environment operational discipline.

  • Excellent written and verbal communication skills; strong documentation habits (knowledge base, runbooks, diagrams, procedures).

  • Ability to work under deadlines, switch contexts quickly, and deliver across multiple initiatives.



Additional Information

Nice-to-Haves

  • Experience supporting regulated healthcare or PHI-adjacent environments and governance expectations.

  • Experience supporting SOC2/ISO-style audits (evidence, control operation, policy-driven operations).

  • Familiarity with internal IT tooling and identity/access systems (e.g., SSO, VPN, device management patterns).

  • Experience building internal developer platforms or “golden path” delivery tooling.

Create a job alert for this search

Lead Platform Engineer (DevOps & MLOps) • Toronto, Ontario, Canada

Similar jobs

Solution Developer, Power Platform & D365

Debtt GroupMarkham, ON, CA
Permanent

Solution Developer, Power Platform & D365.Mentions agentic / vibe coding and Copilot Studio; uses AI to augment development, design and testing workflows.Join MNP as a Solution Developer focused on... Show more

 • Promoted

Remote Mid-Level Platform Engineer

NTT DATA North AmericaToronto, ON, CA
Remote
Full-time

If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Mid-Level Platform Engineer to join our team remotely in Canada.NTT DATA’s... Show more

 • Promoted

Remote Platform Engineer — Cloud & Kubernetes Ops

PlanetToronto, ON, CA
Remote
Full-time

A leading global space and data company is seeking a Software Engineer in Platform Operations.This full-time remote role prioritizes building and operating cloud infrastructure supporting engineeri... Show more

 • Promoted

Forward Deployed Engineer

DoppelToronto, Ontario, Canada
Full-time

Doppel is building the future of social engineering defense.Our AI-native platform uses agentic AI to protect executives, employees, customers, and brands from phishing, impersonation, fraud, and o... Show more

 • Promoted

Senior Platform Engineer: Hybrid Cloud & Kubernetes Leader

ScotiabankToronto, Canada
Full-time

A leading financial institution in Toronto is seeking a Senior Platform Engineer to lead the design and delivery of hybrid cloud platforms.The successful candidate will have extensive experience wi... Show more

 • Promoted

Cloud Platform Engineer — Aws, Kubernetes & Devops

CI FinancialToronto, Canada
Full-time

A leading financial services firm in Toronto is looking for a Cloud Platform Engineer to work within a cross-functional team.This role involves designing and implementing robust cloud solutions usi... Show more

 • Promoted

Platform Lead

Q1 Technologies, Inc.Toronto, Canada
Full-time

As a Platform Lead for our Core Banking team, you oversee the technical direction, development, and implementation of our technology platforms within the organization.This role requires a deep unde... Show more

 • Promoted

DevOps Engineer

PaymentusRichmond Hill
Full-time

The DevOps Engineer is responsible for supporting, monitoring and tooling of cloud deployments.This engineer works closely with the Development and QA teams to produce reliable and secure productio... Show more

 • Promoted • New!

Senior Ml Engineer — Mlops & Cloud Ai Platform - C$110,000 - C$145,000 A Year

Aviva plcEast York, Canada
Full-time

Designs and deploys ML solutions, focusing on MLOps and cloud AI platforms.Requires Python and AWS proficiency. Show more

 • Promoted

Platform Engineer - C$140,000 - C$250,000 A Year - Remote

RealmEast York, Canada
Remote
Full-time

Seeking a Platform Engineer to design, build, and maintain a data platform.Responsibilities include cloud-native services, infrastructure as code (Kubernetes, Pulumi, Terraform), CI/CD, and observa... Show more

 • Promoted

Senior Platform Engineer – Cloud Infrastructure & Devops - $150,000 - $200,000 A Year

QuickplayEast York, Canada
Full-time

Design and build CI/CD pipelines, manage infrastructure as code, and support microservice containerization, using tools like Terraform, Kubernetes, and Docker, focusing on cloud migration and stake... Show more

 • Promoted

ML Platform Engineering Team – DevOps engineer

Themesoft Inc.Toronto
Full-time +2

Be among the first 25 applicants.Get AI-powered advice on this job and more exclusive features.Direct message the job poster from Themesoft Inc.IT solutions provider and a Woman‑Owned Minority Busi... Show more

 • Promoted • New!

Senior DevSecOps Engineer - Platform Modernization

ApexonToronto, ON, CA
Full-time

Lead platform modernization as a Senior DevSecOps Engineer with a significant financial institution.Bring your 10+ years of software engineering expertise to drive automation and engineering transf... Show more

 • Promoted

Mid-Level DevOps Engineer for Cloud and Embedded Software

ZRG CareersMarkham, York region, Canada
Full-time

Exciting hybrid role for a savvy DevOps Engineer to implement and improve CI/CD pipelines for cloud-based and embedded software solutions.Contribute to optimizing development environments and deliv... Show more

 • Promoted

Lead Debug Engineer for Datacenter GPUs

Advanced Micro DevicesMarkham, Ontario, Canada
Full-time

Join AMD as a Lead Debug Engineer in Markham, Ontario, focused on advanced hardware validation for Datacenter GPU technologies.Leverage your expertise in signal integrity and problem-solving.As a S... Show more

 • Promoted

Ml Platform Engineering Team – Devops Engineer - C$65 - C$75 An Hour

Themesoft IncNorth York, Canada
Full-time

DevOps Engineer to deploy and modernize Machine Learning Kubernetes infrastructure, ensuring deployments comply with enterprise security standards. Show more

 • Promoted

Cloud & DevOps Platform Leader (Hybrid, Toronto)

Targeted TalentToronto, ON, CA
Full-time

A technology solutions provider in Toronto is seeking a Manager for Cloud & DevOps Platform Support.This role involves overseeing a team, providing technical support, and promoting modern Agile pra... Show more

 • Promoted

Platform Engineer

Fiat RepublicToronto, Canada
Full-time

About Fiat RepublicFiat Republic is a London-based, remote-friendly fintech dedicated to bringing mainstream banking to the world of digital assets.Founded in 2021, the company has quickly establis... Show more

 • Promoted

Senior DevOps Engineer for Azure Solutions

LeadingtalentMarkham, ON, CA
Full-time

Elevate your career as a Senior DevOps Engineer for Azure Solutions at IBM, where you'll enhance HashiCorp’s Terraform offerings for multi-cloud environments.This role emphasizes security, automati... Show more

 • Promoted

Azure DevOps Engineer

SwoonMarkham, York Region, CA
Full-time

This range is provided by Swoon.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Direct message the job poster from Swoon.Technical Recruiter @ S... Show more