Talent.com
TOTEM Recruteur de talent
Site Reliability EngineerTOTEM Recruteur de talent • Toronto, ON, CA
Site Reliability Engineer

Site Reliability Engineer

TOTEM Recruteur de talent • Toronto, ON, CA
1 day ago
Job type
  • Permanent
Job description

Employment Status: Permanent
Schedule: 40 hours/week – 100% remote work



Job Description

We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms.

Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-based services. This role combines hands-on engineering with operational leadership, giving you direct ownership of system availability, scalability, and incident response.

You will be involved throughout the service lifecycle, from architecture and launch preparation to production monitoring and continuous improvement.


Responsibilities

  • Partner with engineering teams during system design, capacity planning, launch readiness, and production deployment.
  • Monitor and improve service availability, latency, performance, and overall system health.
  • Identify recurring operational issues and implement sustainable solutions that improve scalability and resilience.
  • Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs.
  • Build and maintain automated infrastructure using Terraform and CI/CD pipelines.
  • Develop automation and operational tooling to reduce manual intervention and support self-healing systems.
  • Coordinate incident response and act as Incident Commander during critical production events.
  • Facilitate blameless post-incident reviews and ensure that corrective actions are completed.
  • Use AI-assisted engineering tools responsibly to accelerate development and operational workflows.
  • Maintain clear technical documentation and contribute to the continuous improvement of SRE practices.


Required profile

  • Significant experience in Site Reliability Engineering, Cloud Engineering, DevOps, or a similar infrastructure-focused role.
  • Experience supporting complex or large-scale SaaS environments with high availability requirements.
  • Strong hands-on knowledge of AWS services and architecture, including multi-account environments, VPC, EC2, and EKS.
  • Proven experience operating and troubleshooting Kubernetes environments at scale.
  • Strong knowledge of Infrastructure as Code, particularly Terraform.
  • Experience building or maintaining CI/CD pipelines using GitLab, Jenkins, or comparable tools.
  • Experience with enterprise observability platforms such as Datadog, Prometheus, Grafana, or equivalent solutions.
  • Strong scripting skills using Python, Bash, or a similar language.
  • Direct experience participating in on-call rotations, coordinating incident response, and conducting post-incident reviews.
  • Familiarity with Java or .NET application environments is considered an asset.
  • Ability to communicate clearly and collaborate with development, infrastructure, security, and operations teams.
  • Must be legally authorized to work in Canada.


What to Expect

  • Remote-first work environment within Canada.
  • Occasional visits to a local office or participation in in-person meetings may be required, representing less than 10% of the role.
  • Participation in a scheduled on-call rotation is required.


Does this opportunity sound like a good fit for you? Apply now through our website or by sending your resume to e.henry@totemtalent.ca.

Thank you for your interest in this position; only candidates who meet our client’s requirements will be contacted.

The masculine gender is used as a neutral form.

#totemtech

Create a job alert for this search

Site Reliability Engineer • Toronto, ON, CA

Similar jobs

Site Reliability Engineer, Observability - C$110,000 - C$130,000 A Year

PricelineNorth York, Canada
Full-time

Seeking a Site Reliability Engineer with 3+ years of experience in Observability, SRE, or DevOps to enhance production visibility, improve system reliability, and support a global scaling environment. Show more

 • Promoted

Site Reliability Engineer Role at Magnet Forensics

Magnet ForensicsToronto, ON, CA
Full-time

Take your expertise to the next level as a Senior Site Reliability Engineer with Magnet Forensics.This role involves hands-on AWS and Kubernetes management to uphold our SaaS platform’s reliability... Show more

 • Promoted

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Full-time

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, Ontario, Canada
Full-time

Working Location: Toronto, ON (Hybrid 2 days a week in office).The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Show more

 • Promoted

Senior Site Reliability Engineer

Guidewire SoftwareToronto, Ontario, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Site Reliability Engineer

KyndrylToronto, ON, CA
Full-time +1

Join to apply for the Site Reliability Engineer role at Kyndryl.Direct message the job poster from Kyndryl.Recruitment & Strategic Staffing @Kyndryl | Partnering with IT Consultants in Financial Se... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Senior Site Reliability Engineer, Kong Konnect

Kong Inc.Toronto, ON, CA
Full-time

Senior Site Reliability Engineer, Kong Konnect.Join to apply for the Senior Site Reliability Engineer, Kong Konnect role at Kong Inc.Are you ready to power the World's connections?.If you don’t thi... Show more

 • Promoted

Site Reliability Engineer - C$113,400 - C$162,000 A Year

TextNowNorth York, Canada
Full-time

Senior Site Reliability Engineer to design and maintain scalable systems, automate infrastructure, and improve system reliability. Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

IBM Site Reliability Engineering Expert

LeadingtalentMarkham
Full-time

Step into a career as a Site Reliability Engineer at IBM, focused on enhancing system reliability and performance.Engage directly with production systems and optimize customer experience.In this ro... Show more

 • Promoted

Senior Site Reliability Engineer

Sage Recruiting Inc.Toronto, Ontario, Canada
Full-time

This range is provided by Sage Recruiting Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Senior Site Reliability Engineer (Founding Role).A... Show more

 • Promoted

Site Reliability Engineer Iii - C$100,000 - C$126,000 A Year

AcvNorth York, Canada
Full-time

The SRE will build and ship new features, optimize operational efficiency, and drive growth. Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto
Full-time +1

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Site Reliability Engineer - C$102,700 - C$137,000 A Year

McCain FoodsNorth York, Canada
Full-time

Seeking a Site Reliability Engineer to ensure software system reliability and availability through resilient architecture, automation, and incident response.Responsibilities include system design i... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer

Momentum Financial Services GroupToronto
Full-time

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more