Talent.com

Reliability Jobs in Toronto, ON

Create a job alert for this search

Reliability • toronto on

Last updated: 3 days ago

SRE – Technical Project Manager || Toronto, ON - Onsite || Fulltime FTE

AcestackToronto, ON, Canada
Full-time
Quick Apply

SRE Technical Project Manager</b></div> <div><b>Location: Toronto, ON<br /> Work Arrangement: Onsite<br /> Employment Type: Full-Time FTE</b></div&g... Show more

Reliability Engineer

Kinross Gold CorporationToronto, ON, CA
Full-time

Location: Downtown Toronto (outside Union Station – TTC & GO accessible).Founded in 1993, Kinross is a Canadian-based senior gold mining company with operations and projects in the United State... Show more

Site Reliability Engineer - SRE

Royal Bank of Canada>TORONTO, Canada
Full-time

We are seeking a Site Resiliency Engineer to be a member of Production Support team and focus on automation, development, implementation and support of Site Reliability Engineering (SRE) solutions.... Show more

Maintenance Repair Analyst

Seaboard Transport GroupNorth York, ON, CA
Full-time

The Maintenance Repair Analyst plays a key role in strengthening fleet reliability, safety and performance for heavy-duty trucks and trailers.By connecting field expertise with fleet leadership, th... Show more

Technical Co-founder (CTO) - AI Legal Front Desk

FutureSightToronto, ON, CA
Remote
Full-time
Quick Apply

Solo practitioners and small law firms still lose work the same way: missed calls, slow callbacks, and intake handled between hearings.Hiring is no answer — payroll, training, and turnover are exac... Show more

Site Reliability Engineer- TDJP00058343

Randstad CanadaToronto, Ontario, CA
Full-time +2
Quick Apply

Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team.In this engineering-focus... Show more

Manager, Network Reliability and Resiliency

ServiceNowToronto, Ontario, Canada
CA$125,700.00 yearly
Full-time +1

Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition ... Show more

Junior RAMS/Systems Assurance Engineer

Egis GroupToronto, Ontario, Canada
Full-time

The RAMS Engineer supports the development and execution of the Reliability, Availability, Maintainability, and Safety (RAMS) program for the LRT project.This role is responsible for conducting RAM... Show more

Site Reliability Engineer

TOTEM Recruteur de talentToronto, ON, CA
Permanent

Schedule: 40 hours/week – 100% remote work.We are looking for an experienced.Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infr... Show more

Senior Site Reliability Engineer

iManageToronto, ON, CA
Full-time
Quick Apply

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams -- SRE teams are anchored ... Show more

Sr. Builder - Mobile (Sr. SDE), Ring

Amazon Development Centre Canada ULC - K03Toronto, Ontario, CAN
Full-time

Ring is redefining how millions of people interact with their homes every single day.As a Senior Builder on this feature team, you'll own and evolve some of the most foundational user experiences i... Show more

Microsoft Power Platform Engineer - Canada Only

Blue MantisToronto, Ontario, CA
Full-time
Quick Apply

You Must Be Located In Canada .The Power Platform Engineer is accountable for designing, implementing, and supporting Blue Mantis’ Power Platform solutions to enable scalable, secure, and enterpris... Show more

Mechanical Engineer – Onshore Reliability

Hudson ManpowerToronto, ON, CA
Full-time

Mechanical Engineer – Onshore Reliability.Bachelor’s Degree in Mechanical Engineering.Oil & Gas / Refinery (Onshore).The Mechanical Engineer – Onshore Reliability will be responsible for improv... Show more

Software Development Engineer, Personnel Resource Manager

Amazon Development Centre Canada ULCToronto, Ontario, CAN
Full-time

The Worblehat team is looking for a Software Development Engineer to join our mission-critical platform that automates operational compliance across Amazon's global fulfillment network.This unique ... Show more

Executive Director, Reliability & Maintenance Services

ConfidentialToronto, Ontario, CA
Full-time

Executive Director, Reliability & Maintenance Services About the Company Well-regarded organization managing a local airport Industry Airlines/Aviation Type Non Profit Founded 1996 Employees ... Show more

 • Promoted

Reliability Expert - Fully Remote | Upto $120/hr

MercorToronto, Ontario, Canada
CA$80.00 hourly
Remote
Part-time
Quick Apply

Headquartered in San Francisco, our investors include.Incident management / reliability / SRE Evaluator.Evaluate AI-generated artifacts against domain-specific quality rubrics.Identify factual, aes... Show more

Sr. Machine Learning Software Verification Engineer

TalentlabToronto, Ontario, Canada
Full-time

AI Software Test / Validation Engineer.Technology / AI / Semiconductor.Our client is a global technology leader developing next-generation AI and machine learning solutions for on-device applicatio... Show more

Mid-Senior Mining Professionals

Hire Resolve.comToronto, ON, CA
Full-time
Quick Apply

Hire Resolve is assisting mining organizations in hiring experienced mining professionals across Canada.This is a multi-role opportunity spanning several functions within the sector, including mine... Show more

AI Infrastructure Engineer

Palona AIToronto, ON, CA
Full-time
Quick Apply

Palona’s AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks.Infrastructure is therefore part of the p... Show more

People also ask
SRE – Technical Project Manager || Toronto, ON - Onsite || Fulltime FTE

SRE – Technical Project Manager || Toronto, ON - Onsite || Fulltime FTE

AcestackToronto, ON, Canada
3 days ago
Job type
  • Full-time
  • Quick Apply
Job description

SRE Technical Project Manager

Location: Toronto, ON
Work Arrangement: Onsite
Employment Type: Full-Time FTE

Job Summary: We are seeking an experienced SRE / DevOps Engineer Technical Project Manager focused on reliability engineering, automation, observability, and cloud operations. The ideal candidate will have strong hands-on expertise with Dynatrace, AWS, Azure, Ansible, Terraform, CI/CD, and Kubernetes, along with the ability to coordinate technical initiatives and drive reliability improvements.

Required Skills & Qualifications

  • Strong expertise in Dynatrace, including APM, Davis AI, RUM, and infrastructure monitoring.
  • Experience using Dynatrace Davis AI for root cause analysis, anomaly detection, predictive insights, and alert optimization.
  • Hands-on experience with OneAgent, Smartscape, distributed tracing, SLIs/SLOs, dashboards, and alert management.
  • Strong automation experience using Ansible for deployments, provisioning, patching, and remediation.
  • Hands-on cloud experience with AWS (primary) and Azure, including serverless and cloud-native architectures.
  • Strong Infrastructure as Code (IaC) experience with Terraform, CloudFormation, or AWS CDK.
  • Experience implementing CI/CD pipelines using Jenkins, GitHub Actions, and GitLab CI.
  • Experience integrating monitoring and observability tools with CI/CD pipelines.
  • Strong knowledge of Docker, Kubernetes, ECS, and AKS.
  • Experience with High Availability (HA), Disaster Recovery (DR), incident response, and reliability engineering.
  • Strong programming/scripting skills in Python (boto3) and Bash.
  • Experience with AWS CloudWatch and Azure Monitor.
  • Exposure to Prometheus, Grafana, and ELK is an advantage.

Key Responsibilities

  • Design, implement, and maintain reliable, scalable, and highly available cloud infrastructure.
  • Lead observability initiatives using Dynatrace across applications, infrastructure, and user experience.
  • Configure and optimize Dynatrace OneAgent, Smartscape, distributed tracing, dashboards, alerts, SLIs, and SLOs.
  • Leverage Davis AI for automated anomaly detection, root cause analysis, predictive insights, and alert optimization.
  • Develop and maintain Ansible playbooks for deployment, provisioning, patching, and automated remediation.
  • Automate infrastructure provisioning and configuration using Terraform, CloudFormation, or CDK.
  • Build and maintain CI/CD pipelines and integrate observability and monitoring capabilities into deployment workflows.
  • Support containerized workloads using Docker, Kubernetes, ECS, and AKS.
  • Implement and maintain HA, DR, monitoring, alerting, and incident response processes.
  • Troubleshoot complex application, infrastructure, cloud, and network reliability issues.
  • Collaborate with development, infrastructure, security, cloud, and business teams to improve system reliability and operational efficiency.
  • Drive automation and continuous improvement across SRE and DevOps processes.
  • Provide technical leadership and coordinate delivery of reliability, observability, and automation initiatives.

Preferred Qualifications

  • Experience in Site Reliability Engineering (SRE), DevOps, Cloud Engineering, or Technical Project Management.
  • Strong understanding of cloud-native architecture and enterprise observability.
  • Excellent communication, stakeholder management, problem-solving, and technical leadership skills.
  • Experience managing multiple technical initiatives in a fast-paced enterprise environment.