Talent.com
Axelon Services Corporation
Site Reliability EngineerAxelon Services Corporation • Montreal, QC
No longer accepting applications
Site Reliability Engineer

Site Reliability Engineer

Axelon Services Corporation • Montreal, QC
24 days ago
Job type
  • Temporary
Job description

Summary:

  • Duration: 12 Months Contract
  • Work Mode: Full Remote

Responsibilities:

  • Maintain and improve the reliability, availability, scalability, and performance of cybersecurity platforms, services, and supporting infrastructure.
  • Support day-to-day operational stability by monitoring system health, identifying risks, responding to incidents, and driving timely resolution of service-impacting issues.
  • Instrument infrastructure, applications, services, APIs, data pipelines, and cloud components to provide end-to-end visibility into system behavior and service health.
  • Design, build, and continuously refine monitoring, alerting, logging, tracing, and observability capabilities across distributed systems and cloud environments.
  • Develop meaningful and actionable alerts that reduce noise, improve signal quality, and enable teams to respond quickly to emerging issues.
  • Define and track key reliability metrics, including availability, latency, throughput, error rates, saturation, service-level indicators, service-level objectives, and operational risk indicators.
  • Build, maintain, and enhance dashboards for engineering, operations, product, risk, and executive stakeholders, ensuring information is accurate, timely, and decision-ready.
  • Continuously modify and improve executive dashboards to support regular leadership reviews of service health, reliability trends, incidents, risks, and operational performance.
  • Partner with engineering, cybersecurity, infrastructure, cloud, and application teams to identify reliability gaps and implement long-term improvements.
  • Participate in incident response, root-cause analysis, problem management, and post-incident reviews to prevent recurrence and improve operational maturity.
  • Automate operational tasks, health checks, reporting, deployment validation, and recovery procedures to improve efficiency and reduce manual effort.
  • Collaborate with application and platform teams to embed reliability, monitoring, and supportability requirements into the software development lifecycle.
  • Support CI/CD, DevOps, and release management practices by validating operational readiness, monitoring coverage, rollback plans, and production support requirements.
  • Contribute to resiliency engineering efforts, including capacity planning, performance tuning, failover validation, disaster recovery readiness, and chaos/resilience testing where applicable.
  • Ensure monitoring, alerting, dashboards, and operational processes align with enterprise security, risk, compliance, and governance standards.

Requirements:

  • Minimum 10 years of experience in site reliability engineering, systems engineering, software engineering, DevOps, infrastructure engineering, or production operations.
  • Strong experience supporting highly available, distributed, cloud-based, or mission-critical technology platforms.
  • Hands-on experience with observability practices, including monitoring, alerting, logging, metrics, tracing, dashboards, and service health reporting.
  • Experience instrumenting applications, services, APIs, infrastructure, databases, and cloud components to enable end-to-end operational visibility.
  • Strong understanding of reliability engineering concepts, including SLIs, SLOs, SLAs, error budgets, incident management, capacity management, and operational readiness.
  • Experience designing actionable alerts that support rapid issue detection, triage, escalation, and resolution.
  • Experience building and maintaining operational dashboards for technical teams, support teams, and senior/executive stakeholders.
  • Strong scripting or programming skills using Python, Java, Bash, PowerShell, or similar languages for automation and operational tooling.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Experience with Infrastructure-as-Code tools such as Terraform or similar technologies.
  • Experience working with CI/CD pipelines, DevOps workflows, release processes, and production support models.
  • Experience troubleshooting distributed systems, REST services, event-driven architectures, messaging platforms, and service-to-service integrations.
  • Familiarity with relational and non-relational databases, such as PostgreSQL, MSSQL, MongoDB, or similar platforms.
  • Strong analytical, troubleshooting, and problem-solving skills with the ability to diagnose complex technical issues across multiple layers of the stack.
  • Strong written and verbal communication skills, including the ability to translate technical issues into clear business and executive-level updates.

Preferred Skills:

  • Experience supporting cybersecurity, risk, resilience, compliance, or enterprise security platforms.
  • Experience with observability and monitoring tools such as Splunk, Grafana, Prometheus, Datadog, Dynatrace, New Relic, Azure Monitor, CloudWatch, OpenTelemetry, or similar platforms.
  • Experience creating executive-level service health dashboards, reliability scorecards, operational risk reporting, or incident trend reporting.
  • Experience developing automated health checks, synthetic monitoring, service dependency maps, and operational runbooks.
  • Experience with incident response, major incident management, postmortems, root-cause analysis, and problem management practices.
  • Experience with containerized and cloud-native environments, including Kubernetes, Docker, serverless services, or managed cloud platforms.
  • Experience with distributed messaging or streaming platforms such as Apache Kafka.
  • Familiarity with cloud-native security, governance, and policy tooling such as Azure Policy, AWS SCP, GCP constraints, or related controls.
  • Familiarity with Cloud Security Posture Management tools such as Wiz, Prisma, CloudGuard, or similar platforms.
  • Experience with cloud-based AI services such as Azure AI, AWS Bedrock, or Google Vertex AI, particularly from an operational monitoring, reliability, or governance perspective.
  • Experience supporting Linux and Windows environments through scripting, automation, monitoring, and operational troubleshooting.
  • Exposure to web technologies, APIs, front-end services, or user-facing application monitoring.

Additional Skills:

  • Strong ownership mindset with a focus on operational excellence and service reliability.
  • Ability to operate effectively in fast-paced, production-focused environments with minimal supervision.
  • Strong ability to prioritize issues based on customer impact, business risk, service criticality, and operational urgency.
  • Effective collaboration skills across engineering, operations, cybersecurity, infrastructure, risk, and executive stakeholder groups.
  • Ability to communicate service health, operational risks, incidents, and reliability trends clearly to both technical and non-technical audiences.
  • Proactive and continuous-improvement mindset with a focus on automation, simplification, resilience, and measurable outcomes.
  • Strong attention to detail when building dashboards, defining metrics, tuning alerts, and preparing executive-level operational reporting.
Create a job alert for this search

Site Reliability Engineer • Montreal, QC

Similar jobs

Site Reliability Engineer

Vertex Elite LLCRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Duration: Contract Key Skills: Monitoring / Observability tools - Dynatrace, ELK etc.Platform/ cloud Observability - OpenShift, Prometheus / Azure Cloud etc.Key Responsibilities: Collaborate with v... Show more

 • Promoted

Senior Platform Engineer - Remote, Scale & Reliability

Lillio (formerly HiMama)Montreal (administrative region), QC, CA
Remote
Full-time

A leading EdTech company in Canada is seeking a Senior Platform Engineer to enhance system reliability and performance while contributing to impactful software solutions.The role involves making te... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsMontreal (administrative region), QC, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Site Reliability Engineer

ApTaskMontréal, Canada
Full-time

Direct message the job poster from ApTask Looking for an intermediate between 2 to 5 years' experience.The Application Infrastructure (Al) department is seeking a Site Reliability Engineer (SRE... Show more

 • Promoted

REMOTE Protection & Controls Substation Engineer

JobotMontreal (administrative region), QC, CA
Remote
Full-time

REMOTE Protection & Controls Substation Engineer.Be among the first 25 applicants.This range is provided by Jobot.Your actual pay will be based on your skills and experience — talk with your recrui... Show more

 • Promoted

Production Systems Administrator: Automation & Reliability

UbisoftMontreal (administrative region), QC, CA
Full-time

A global gaming leader in Montreal is seeking a System Administrator to ensure the operational stability of infrastructure supporting game production.In this role, you'll design automation scripts,... Show more

 • Promoted

Remote Engineering Team Lead - Grow Engineers and Impact

RemoteMontreal (administrative region), QC, CA
Remote
Full-time

A global employment solutions company is seeking a Team Leader to manage and guide a small engineering product team.Responsibilities include overseeing daily operations, engaging in hiring and onbo... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificMontreal (administrative region), QC, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Senior Contaminated Sites Lead - Remote-Eligible

Stantec Consulting International Ltd.Montreal (administrative region), QC, CA
Remote
Full-time

A global engineering firm seeks a Senior Environmental Professional to lead environmental investigations and remediation projects across Canada.The role involves project management, client relation... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Infrastructure Reliability Engineer

ShippoMontreal (administrative region), QC, CA
Full-time

Enhance shipping solutions as a Senior Site Reliability Engineer in a remote setting.Focus on infrastructure integrity, scalability, and performance in a collaborative environment.This position inv... Show more

 • Promoted

Senior Full Stack Engineer in Remote Agile Development Team

BevertecMontreal (administrative region), QC, CA
Remote
Full-time

Become a part of a remote Agile team as a Senior Full Stack Engineer.Focus on developing applications that enhance community information sharing and support funding processes in Canada.This positio... Show more

 • Promoted

Site Supervisor

Nasittuq CorporationMontreal (administrative region), QC, CA
Full-time

Join Nasittuq for a unique and rewarding experience!.Nasittuq Corporation (from the Inuktitut word meaning “looking out from the highest point”) operates and maintains the North Warning System (NWS... Show more

 • Promoted

Experienced Site Reliability Engineer - Remote

Tech InsightsMontreal (administrative region), QC, CA
Remote
Full-time

TechInsights seeks a Senior Site Reliability Engineer to enhance AI operations from anywhere in Canada.Oversee reliability strategies, manage error budgets, and collaborate closely with engineering... Show more

 • Promoted

Remote Site Reliability Engineer - Scale Crypto Systems

NewtonMontreal (administrative region), QC, CA
Remote
Full-time

A leading innovative tech company in Toronto is looking for a Site Reliability Engineer.In this pivotal role, you will enhance the reliability and resilience of critical services, manage incidents,... Show more

 • Promoted

Staff Site Reliability Engineer, Database

AlpacaMontreal (administrative region), QC, CA
Full-time

Alpaca is a US-headquartered self-clearing broker-dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.Our recent Series D funding round broug... Show more

 • Promoted

Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT GroupMontréal, Quebec, Canada
Full-time

Site Reliability Engineer (Linux / Cloud Infrastructure) role with hands-on experience across Linux, distributed systems, scripting, databases, monitoring, containers, cloud SaaS integrations, mess... Show more

 • Promoted

Site Quality Assurance Engineer for Hydropower

Andritz AGPointe-Claire, Montreal (administrative region), CA
Full-time

Ensure top-notch quality in hydropower installations as a Site Quality Assurance Engineer.Dominate quality control processes and specifications while working at various project locations in Canada.... Show more