Talent.com
Axelon Services Corporation
Site Reliability EngineerAxelon Services Corporation • Toronto, ON
Site Reliability Engineer

Site Reliability Engineer

Axelon Services Corporation • Toronto, ON
4 days ago
Job type
  • Temporary
Job description

Summary:

  • Location: Remote across Canada
  • Duration: 12 months contract

Responsibilities:

  • Maintain and improve the reliability, availability, scalability, and performance of cybersecurity platforms, services, and supporting infrastructure.
  • Support day-to-day operational stability by monitoring system health, identifying risks, responding to incidents, and driving timely resolution of service-impacting issues.
  • Instrument infrastructure, applications, services, APIs, data pipelines, and cloud components to provide end-to-end visibility into system behavior and service health.
  • Design, build, and continuously refine monitoring, alerting, logging, tracing, and observability capabilities across distributed systems and cloud environments.
  • Develop meaningful and actionable alerts that reduce noise, improve signal quality, and enable teams to respond quickly to emerging issues.
  • Define and track key reliability metrics, including availability, latency, throughput, error rates, saturation, service-level indicators, service-level objectives, and operational risk indicators.
  • Build, maintain, and enhance dashboards for engineering, operations, product, risk, and executive stakeholders, ensuring information is accurate, timely, and decision-ready.
  • Continuously modify and improve executive dashboards to support regular leadership reviews of service health, reliability trends, incidents, risks, and operational performance.
  • Partner with engineering, cybersecurity, infrastructure, cloud, and application teams to identify reliability gaps and implement long-term improvements.
  • Participate in incident response, root-cause analysis, problem management, and post-incident reviews to prevent recurrence and improve operational maturity.
  • Automate operational tasks, health checks, reporting, deployment validation, and recovery procedures to improve efficiency and reduce manual effort.
  • Collaborate with application and platform teams to embed reliability, monitoring, and supportability requirements into the software development lifecycle.
  • Support CI/CD, DevOps, and release management practices by validating operational readiness, monitoring coverage, rollback plans, and production support requirements.
  • Contribute to resiliency engineering efforts, including capacity planning, performance tuning, failover validation, disaster recovery readiness, and chaos/resilience testing where applicable.
  • Ensure monitoring, alerting, dashboards, and operational processes align with enterprise security, risk, compliance, and governance standards.

Requirements:

  • Minimum 10 years of experience in site reliability engineering, systems engineering, software engineering, DevOps, infrastructure engineering, or production operations.
  • Strong experience supporting highly available, distributed, cloud-based, or mission-critical technology platforms.
  • Hands-on experience with observability practices, including monitoring, alerting, logging, metrics, tracing, dashboards, and service health reporting.
  • Experience instrumenting applications, services, APIs, infrastructure, databases, and cloud components to enable end-to-end operational visibility.
  • Strong understanding of reliability engineering concepts, including SLIs, SLOs, SLAs, error budgets, incident management, capacity management, and operational readiness.
  • Experience designing actionable alerts that support rapid issue detection, triage, escalation, and resolution.
  • Experience building and maintaining operational dashboards for technical teams, support teams, and senior/executive stakeholders.
  • Strong scripting or programming skills using Python, Java, Bash, PowerShell, or similar languages for automation and operational tooling.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience with Infrastructure-as-Code tools such as Terraform or similar technologies.
  • Experience working with CI/CD pipelines, DevOps workflows, release processes, and production support models.
  • Experience troubleshooting distributed systems, REST services, event-driven architectures, messaging platforms, and service-to-service integrations.
  • Familiarity with relational and non-relational databases, such as PostgreSQL, MSSQL, MongoDB, or similar platforms.
  • Strong analytical, troubleshooting, and problem-solving skills with the ability to diagnose complex technical issues across multiple layers of the stack.
  • Strong written and verbal communication skills, including the ability to translate technical issues into clear business and executive-level updates.

Preferred Skills:

  • Experience supporting cybersecurity, risk, resilience, compliance, or enterprise security platforms.
  • Experience with observability and monitoring tools such as Splunk, Grafana, Prometheus, Datadog, Dynatrace, New Relic, Azure Monitor, CloudWatch, OpenTelemetry, or similar platforms.
  • Experience creating executive-level service health dashboards, reliability scorecards, operational risk reporting, or incident trend reporting.
  • Experience developing automated health checks, synthetic monitoring, service dependency maps, and operational runbooks.
  • Experience with incident response, major incident management, postmortems, root-cause analysis, and problem management practices.
  • Experience with containerized and cloud-native environments, including Kubernetes, Docker, serverless services, or managed cloud platforms.
  • Experience with distributed messaging or streaming platforms such as Apache Kafka.
  • Familiarity with cloud-native security, governance, and policy tooling such as Azure Policy, AWS SCP, GCP constraints, or related controls.
  • Familiarity with Cloud Security Posture Management tools such as Wiz, Prisma, CloudGuard, or similar platforms.
  • Experience with cloud-based AI services such as Azure AI, AWS Bedrock, or Google Vertex AI, particularly from an operational monitoring, reliability, or governance perspective.
  • Experience supporting Linux and Windows environments through scripting, automation, monitoring, and operational troubleshooting.
  • Exposure to web technologies, APIs, front-end services, or user-facing application monitoring.

This role is for an existing vacancy.

Create a job alert for this search

Site Reliability Engineer • Toronto, ON

Similar jobs

Senior Site Reliability Engineer

Guidewire Softwaretoronto, on, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Senior Site Reliability Engineer

RelayToronto, ON, CA
Full-time

Relay is a digital banking platform that gives self‑made business owners the tools and know‑how to be great with money—bringing clarity, confidence, and control to every dollar earned, so they can ... Show more

 • Promoted

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Full-time

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, Ontario, Canada
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office] Role Mandate.The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the rel... Show more

 • Promoted

Site Reliability Engineer

KyndrylToronto, ON, CA
Full-time +1

Join to apply for the Site Reliability Engineer role at Kyndryl.Direct message the job poster from Kyndryl.Recruitment & Strategic Staffing @Kyndryl | Partnering with IT Consultants in Financial Se... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

IBM Site Reliability Engineering Expert

LeadingtalentMarkham
Full-time

Step into a career as a Site Reliability Engineer at IBM, focused on enhancing system reliability and performance.Engage directly with production systems and optimize customer experience.In this ro... Show more

 • Promoted

Senior Site Reliability Engineer

Sage Recruiting Inc.Toronto, Ontario, Canada
Full-time

This range is provided by Sage Recruiting Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Senior Site Reliability Engineer (Founding Role).A... Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto
Full-time +1

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Senior Site Reliability Engineer - SLS Team

United States Digital Space LLCtoronto, on, Canada
Full-time

Become the backbone of the company’s cloud storage at the forefront of technology as a Senior Site Reliability Engineer.This role offers opportunities in performance tuning and operational safety f... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Dynatrace Expert Site Reliability Engineer

Intelliswift - An LTTS CompanyToronto, ON, CA
Full-time

Become a pivotal Dynatrace Expert as a Site Reliability Engineer.In this hybrid role, you will ensure distributed systems are performant and observable.The position requires a seasoned professional... Show more

 • Promoted

Site Reliability Engineer

Momentum Financial Services GroupToronto
Full-time

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Site Reliability Engineer

PheedLoopToronto, Ontario, Canada
Full-time

Build the tech behind live events.PheedLoop's mission is to help organizers turn ordinary events into unforgettable experiences with event technology that is bold, intuitive, and built to bring peo... Show more