Talent.com

Reliability Jobs in Toronto, ON

Create a job alert for this search

Reliability • toronto on

Last updated: 11 hours ago

Hiring Sr Support Engineer in Montreal, QC

artechToronto, ON
Full-time

Job Title: Sr Support Engineer.Location: Montreal, QC - Hybrid (2-4 Days WFO).Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboar... Show more

 • New!

Director Production Operations

kort paymentsToronto, ON, CA
Full-time
Quick Apply

DIRECTOR PRODUCTION OPERATIONS WHO WE ARE Welcome to KORT Payments, where innovation meets excellence!.We specialize in providing a state-of-the-art omnichannel payments platform designed to make b... Show more

Reliability Engineer

kinross goldToronto, ON, CA
Full-time

Location: Downtown Toronto (outside Union Station – TTC & GO accessible).Founded in 1993, Kinross is a Canadian-based senior gold mining company with operations and projects in the United State... Show more

Director, Global Fraud Technology - Service Reliability & Production Engineering

scotiabankToronto, ON, CA
Full-time

Join a purpose driven winning team, committed to results, in an inclusive and high-performing culture.The Global Fraud Technology team develops and manages enterprise fraud capabilities that protec... Show more

Software Development Engineer, Ring Cloud Connectivity Org

amazon development centre canada ulcToronto, Ontario, CAN
Full-time

Ring is seeking a Mobile Software Development Engineer to join the Cloud Connectivity organization in Toronto, where you'll build and deliver customer-facing mobile experiences that power how milli... Show more

Maintenance Repair Analyst

seaboard transport groupNorth York, ON, CA
Full-time

The Maintenance Repair Analyst plays a key role in strengthening fleet reliability, safety and performance for heavy-duty trucks and trailers.By connecting field expertise with fleet leadership, th... Show more

Manager, Network Reliability and Resiliency

service nowToronto, Ontario, Canada
CA$125,700.00 yearly
Full-time +1

Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment.The screeni... Show more

Systems Reliability Engineer

staffingineToronto, ON, Canada
Full-time
Quick Apply

Job Title: </b><b>Systems Reliability Engineer<br /> Job Location: Toronto, ON<br /> Job Type: Contract</b></p> <p style="text-align:justify; margin-bott... Show more

Site Reliability Engineer

totem recruteur de talentToronto, ON, CA
Permanent

Schedule: 40 hours/week – 100% remote work.We are looking for an experienced.Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infr... Show more

Senior Site Reliability Engineer

i manageToronto, ON, CA
Full-time
Quick Apply

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams -- SRE teams are anchored ... Show more

Sr. Builder - Mobile (Sr. SDE), Ring

amazon development centre canada ulc k03Toronto, Ontario, CAN
Full-time

Ring is redefining how millions of people interact with their homes every single day.As a Senior Builder on this feature team, you'll own and evolve some of the most foundational user experiences i... Show more

Principal AI/ML Engineer

vanguardToronto, ON, CA
Full-time

As a Principal AI Engineer, you will serve as a senior technical leader responsible for transforming state-of-the-art AI research into scalable, production-ready capabilities that create measurable... Show more

Site Reliability Engineer

0000050007 royal bank canadaToronto, Ontario
Full-time

The SRE will be responsible for assisting in the development, implementation and support of Site Reliability Engineering solutions for all applications across a line of business within CNB (City Na... Show more

Site Reliability Engineer – Azure Operations, Databricks Support & Incident Response

astra north infoteckToronto, ON, ca
Full-time

Site Reliability Engineer – Azure Operations, Databricks Support & Incident Response.Waterloo or Toronto - Hybrid (Tuesday, Wednesday and Thursday).As an Intermediate Site Reliability Engineer,... Show more

Site Reliability Engineer- TDJP00058343

randstad canadaToronto, Ontario, CA
Full-time +2
Quick Apply

Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team.In this engineering-focus... Show more

Site Reliability Engineer (SRE)

scotiabankToronto, ON, CA
Full-time

You want to be challenged with complex problem solving taking the learnings forward as continuous improvements.You thrive on supporting critical systems requiring a high level of trust, resilience ... Show more

Reliability Expert - Fully Remote | Upto $120/hr

mercorToronto, Ontario, Canada
CA$80.00 hourly
Remote
Part-time
Quick Apply

Headquartered in San Francisco, our investors include.Incident management / reliability / SRE Evaluator.Evaluate AI-generated artifacts against domain-specific quality rubrics.Identify factual, aes... Show more

Sr. Machine Learning Software Verification Engineer

talentlabToronto, Ontario, Canada
Full-time

AI Software Test / Validation Engineer.Technology / AI / Semiconductor.Our client is a global technology leader developing next-generation AI and machine learning solutions for on-device applicatio... Show more

Mid-Senior Mining Professionals

hire resolve comToronto, ON, CA
Full-time
Quick Apply

Hire Resolve is assisting mining organizations in hiring experienced mining professionals across Canada.This is a multi-role opportunity spanning several functions within the sector, including mine... Show more

AI Infrastructure Engineer

palona aiToronto, ON, CA
Full-time
Quick Apply

Palona’s AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks.Infrastructure is therefore part of the p... Show more

People also ask
Hiring Sr Support Engineer in Montreal, QC

Hiring Sr Support Engineer in Montreal, QC

artechToronto, ON
11 hours ago
Job type
  • Full-time
Job description

Job Title: Sr Support Engineer

Location: Montreal, QC - Hybrid (2-4 Days WFO)

Duration: 06-12 Months

Hourly Pay Rate: 50 CAD/hr

Job Description:

Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboards, and alerting strategies across distributed systems. Drive observability improvements leveraging industry-leading tools to achieve real-time performance insights and comprehensive system visibility. Instrument applications for end-to-end observability, implementing distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures. Troubleshoot complex incidents in production environments, diagnosing root causes across multiple service layers, databases, caches, and APIs under load using SLI/SLO frameworks. Investigate and resolve Azure Kubernetes Service (AKS) infrastructure, ensuring reliability and scalability of containerized workloads with deep proficiency in Terraform and Azure managed services. Translate business requirements into observable, resilient systems that meet defined SLIs/SLOs and drive reliability improvements. Automate operational tasks to reduce toil and improve system resilience through infrastructure-as-code and CI/CD best practices. Lead incident response and remediation for mission-critical systems, conducting blameless postmortems and building resilience through chaos engineering and tabletop exercises. Collaborate cross-functionally with development, platform, and business teams to improve service availability, scalability, and operational excellence.

Required Skills & Qualifications (Must-have qualifications that candidates must meet to be considered)

  • 8 years hands-on experience in observability, SRE, or DevOps roles with proven expertise across infrastructure and application-level reliability.
  • Deep expertise in observability tooling: Dynatrace, ELK, Splunk, and PagerDuty; demonstrated understanding of observability principles (instrumentation, correlation IDs, SLI/SLO frameworks).
  • Advanced proficiency with Azure Kubernetes Service (AKS), Terraform, and Azure managed services; proven ability to design and implement infrastructure-as-code solutions.
  • Strong hands-on experience instrumenting applications for comprehensive observability: distributed tracing, metrics collection, and log aggregation across Node.js and .NET applications.
  • Proven troubleshooting expertise in distributed systems, diagnosing root causes across multiple service layers, databases, caches, and APIs in production environments.
  • Excellent incident management skills: hands-on experience with PagerDuty and ServiceNow; ability to resolve high-severity incidents rapidly and conduct effective root cause analysis.
  • Knowledge of incident, problem, and change management processes, including SRE principles, blameless postmortems, and chaos engineering practices.
  • Exceptional communication and leadership skills.

Preferred Skills & Qualifications (Nice-to-have skills but are not required)

  • Experience with other cloud platforms and services.
  • Familiarity with additional programming languages and frameworks.
  • Experience in financial services or a related industry.

Day-to-Day Responsibilities (key tasks and expectations for the role)

  • Design and deploy observability solutions using Terraform.
  • Drive improvements using industry-leading observability tools.
  • Instrument applications for comprehensive observability.
  • Troubleshoot and resolve complex production incidents.
  • Collaborate with cross-functional teams to enhance service availability and reliability.

For immediate consideration please click APPLY to begin the screening process with Alex.