Talent.com
Open Systems Technologies
Site Reliability Engineer (SRE), ServiceNow, Application InfrastructureOpen Systems Technologies • Montreal, Montreal (administrative region), Canada
Site Reliability Engineer (SRE), ServiceNow, Application Infrastructure

Site Reliability Engineer (SRE), ServiceNow, Application Infrastructure

Open Systems Technologies • Montreal, Montreal (administrative region), Canada
30+ days ago
Job type
  • Full-time
Job description

Site Reliability Engineer (SRE), ServiceNow, Application Infrastructure

2 days ago Be among the first 25 applicants

The Application Infrastructure (AI) department is seeking a Site Reliability Engineer (SRE) to help drive the reliability engineering, operations and customer support services for Morgan Stanley's ServiceNow SaaS implementation. Reporting to a Site Reliability Engineering & Operations Lead.

This role requires delivering a range of SRE practices within a global community of other SREs. This means teaming up with colleagues to deliver reliable, resilient systems without wasteful operational effort.

SRE practices include task optimization and automation, prioritizing technical debt, observability and monitoring dashboards, capacity management, incident response, and problem elimination.

This position specializes in ServiceNow Software as a Service which provides a suite of IT service management capabilities and is integrated with many products such as chatbot technology, on‑call escalation incident management, and a range of other on‑premises infrastructure (including SQL databases, APIs, and web infrastructure). Despite the focus on value‑add development and process delivery, this is also a production‑side, operational role requiring participation in an on‑call rotation from time to time.

Successful candidates for SRE roles in Application Infrastructure have so far come from a variety of backgrounds; maybe a developer today looking to evolve site reliability as a practice, or an infrastructure specialist with an interest in reliability and resilience principles, or a strong system admin who enjoys troubleshooting along with some task automation experience.

Prior experience in the financial services industry is not required, and we welcome candidates from all industries and backgrounds to apply.

Responsibilities

  • Delivery of improvements that will maximize the availability and performance of supported systems through optimized and automated operational tasks, collaborating on the development of operational tools, ongoing problem management, and architecture reviews with colleagues.
  • Troubleshooting ServiceNow issues, and also some on‑premise capabilities in a Linux environment from time to time, collaborating with others to get to the bottom of issues, and agreeing on lasting improvements that can be made.
  • Exploring and delivering observability including metrics, logging, tracing and alerting that can define and measure the target reliability of a product.
  • Being dependable and responsive during agreed hours, like when part of the on‑call rotation with the rest of the global team (with a time‑off in lieu system).
  • A commitment to understanding the Firm's ServiceNow instances and related dependencies, contributing to their documentation.
  • Identification and prioritization of technical debt that can impact client satisfaction or operational efficiency.
  • Give feedback on policy and procedures related to the delivery of SRE and operational practices with a view to continually making the Firm safer and more efficient.

Skills Required

  • The ideal candidate would have at least one of: Software development skills in one or more programming language, e.g. Python, ServiceNow administration or development experience.
  • 7+ years of experience
  • Proficient oral and written communication skills
  • Establishing warm, effective relationships with colleagues to collaborate on successful delivery
  • A dependable team worker with demonstrated commitment to client service
  • Ability to respond appropriately during occasional technical emergencies, like outages.
  • Open to work in on‑call rotation
  • ServiceNow administration or development experience, although this can be acquired by the successful candidate via on‑the‑job learning and training.

Seniority level

Mid‑Senior level

Employment type

Contract

Job function

Staffing and Recruiting

#J-18808-Ljbffr
Create a job alert for this search

Site Reliability Engineer (SRE), ServiceNow, Application Infrastructure • Montreal, Montreal (administrative region), Canada

Similar jobs

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificMontreal (administrative region), QC, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Site Reliability Engineer

Vertex Elite LLCRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Duration: ContractKey Skills:Monitoring / Observability tools - Dynatrace, ELK etc.Platform/ cloud Observability - OpenShift, Prometheus / Azure Cloud etc.Key Responsibilities:Collaborate with vari... Show more

 • Promoted

SRE-DevSecOps Engineer

High Tech GenesisMontreal (administrative region), QC, CA
Full-time

Allowed Staffing Countries: Canada, Costa Rica, Mexico or Brazil, (Remote).High Tech Genesis is seeking a 3-month contractor who can hit the ground running to support our SaaS platform on AWS.Kuber... Show more

 • Promoted

Purolator Lead Systems Engineer Opportunity

Purolator Inc.Montreal (administrative region), QC, CA
Full-time

Shape the future of logistics at Purolator as a Lead Systems Engineer.Drive operational success through team leadership and process optimization at our Montreal Hub.As the optimization lead, you'll... Show more

 • Promoted

Site Reliability Engineer

Hunter BondMontréal, Canada
Full-time

Role: DevOps EngineerClient: Most Elite Tech Firm in CanadaCompensation: Up to $200k CAD + Bonus + PackageLocation: MontrealOverviewAn Elite FinTech Firm is looking for a highly talented DevOps Eng... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsMontreal (administrative region), QC, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Intact Hybrid Site Reliability Engineer

IntactMontreal
Full-time

Join Intact as a Site Reliability Engineer and elevate operational reliability across cloud platforms.This hands-on role employs Azure, AWS, and GCP expertise to manage incidents effectively.The SR... Show more

 • Promoted

Remote ServiceNow DPR Architect & Release Lead

YochanaMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology firm in Canada is looking for a ServiceNow Implementation Specialist to design and implement Digital Product Release (DPR) solutions.The candidate must have over 5 years of Ser... Show more

 • Promoted

Specialist Site Reliability Engineer

Global Talent Alliance, CanadaMontreal
Full-time

About the job Specialist Site Reliability Engineer.The role of the Specialist Site Reliability Engineer (SRE) is to execute RAM analysis and engineering in support of the I&T solutions.The overall ... Show more

 • Promoted

Site Reliability Engineer

BasetenMontréal, Canada
Full-time

About BasetenBaseten powers mission‐critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting applied AI resear... Show more

 • Promoted

Remote ServiceNow Developer – Public Sector Solutions

MindtrisMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology firm is seeking a ServiceNow Developer for a 12-month remote engagement in Canada.This role involves designing, developing, and implementing solutions within ServiceNow, ensuri... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Lead Platform Engineer Enhancing DevOps and System Reliability

Lillio (formerly HiMama)Montreal (administrative region), QC, CA
Full-time

Transform early childhood education as a Senior Platform Engineer focused on system performance and collaborative tooling.Drive key initiatives for scalable, reliable digital platforms.In this pivo... Show more

 • Promoted

Senior Site Reliability Engineer

SecurityScorecardmontreal (administrative region), qc, Canada
Full-time

SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries.Founded in 2013 by security and risk experts Dr.Alex Ya... Show more

 • Promoted

Remote Scrum Master — ServiceNow CSA Expert

Avanciers Inc.Montreal (administrative region), QC, CA
Remote
Full-time

A leading consultancy firm is seeking a skilled Scrum Master with a ServiceNow Certified System Administrator (CSA) certification for a remote role in Canada.The ideal candidate will leverage deep ... Show more

 • Promoted

Site Reliability Engineer

MaintainXMontreal
Full-time

MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 13,000+ customers including Duracell, Shell, Cintas, and Brenntag.We raised $150M in Series D funding ... Show more

 • Promoted

Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT GroupMontréal, Quebec, Canada
Full-time

Site Reliability Engineer (Linux / Cloud Infrastructure) role with hands-on experience across Linux, distributed systems, scripting, databases, monitoring, containers, cloud SaaS integrations, mess... Show more

 • Promoted

Remote Senior ServiceNow Engineer: Platform Ownership

Smart WorkingMontreal (administrative region), QC, CA
Remote
Full-time

A leading remote work company based in Canada is seeking a Senior ServiceNow Engineer to design and maintain solutions across the ServiceNow platform.This role emphasizes ITSM and operational resil... Show more