Talent.com
Tech Talent International
Senior Site Reliability Engineer – Automation & ObservabilityTech Talent International • Montreal, Quebec, Canada
Senior Site Reliability Engineer – Automation & Observability

Senior Site Reliability Engineer – Automation & Observability

Tech Talent International • Montreal, Quebec, Canada
30+ days ago
Salary
CA$110,000.00 yearly
Job type
  • Full-time
  • Permanent
Job description

Tech Talent International (SI) supplies technical talent to a variety of clients ranging from Fortune 100/500/1000 companies to small and mid-sized organizations in Canada/US and Europe.

We currently have a role as aSenior Site Reliability Engineer (SRE) Automation & Observability with our large consulting client working onsite at a major financial services client in the downtown Montreal area

Role: Cybersecurity - Senior Site Reliability Engineer (SRE) Automation & Observability

Type: Permanent or Contract 40 hrs/week

Location: Hybrid - Downtown Montreal QC -(roles starts off 5 days in office for 1st 3 months then turns into hybrid setup 3 days onsite 2 days from home)

Salary: $110000 - $% bonus 3-5 weeks paid vacation RRSP contribution benefits sick/personal days

Position Overview

The Automation team consists of several Subject Matter Experts (SMEs) who assist the Global Process Owner in designing building and maintaining the organizations IT services. While leading the companys IT services team the IT Service Manager strives to develop reliable IT services and improve the organizations existing IT service infrastructure.

IT Service Managers are responsible for maintaining a high standard of service delivery while managing the organizations IT services and anticipating and resolving issues that may arise within company systems or client environments. These services include infrastructure monitoring task automation server asset management and network inventory management.

Change incident problem and request management along with CMDB (Configuration Management Database) functions are core services widely used throughout CIB IT. The ITSM team serves as the bridge between IT and business stakeholders ensuring coordination and predictability for CIB IT and its business operations.

The team includes SMEs focused on key service areas as directed by management with the objective of delivering high-quality services through various platforms that maximize efficiency and consistent results.

Within the Automation & Observability organization the Production Smart Automation team provides production support services for the Analytics Consulting and Digital Assets IT clusters. This includes both functional and technical support as well as project delivery for production and non-production platforms. The team operates globally and consists of approximately 10 members located in Paris Warsaw Mumbai and Montreal.

Key Responsibilities

The Site Reliability Engineer (SRE) will be part of a multidisciplinary team providing Level 1 and Level 2 technical and project support. This is a production-focused role requiring a broad range of technical expertise.

The SRE will work closely with development and infrastructure teams to:

  • Monitor manage and proactively improve the availability and performance of production environments from presentation and application layers through infrastructure layers.
  • Plan and implement application deployments load testing activities and configuration changes.
  • Ensure production environments are operational and available while collaborating with teams to understand user needs.
  • Contribute to medium- and large-scale technical projects including architecture reviews solution design application upgrades and migrations to new platforms.
  • Collaborate on prioritized tasks while providing regular status updates and maintaining focus on target solutions.
  • Understand delivery lifecycle phases to ensure work is completed according to defined specifications and timelines.
  • Identify opportunities to improve operational efficiency and contribute to automation initiatives.
  • Provide constructive feedback and recommendations to management regarding performance capacity and system design.
  • Assist in documenting architectures and designs as well as distributing meeting minutes and action items.

The SRE will also work with other teams to respond to incidents and resolve issues quickly often under pressure in order to restore normal business services. As a result participation in on-call rotations and after-hours support may be required.

Candidates should possess both the aptitude and desire to learn new technologies and contribute innovative ideas that may benefit the department.

Requirements

Candidates should have:

  • 57 years of experience in a similar role.
  • Experience providing multidisciplinary technical support within a team environment.
  • Practical knowledge of performance and capacity management across:
    • Applications
    • Databases
    • Networks
  • Strong automation skills and mindset.

Skills & Competencies

Systems Administration

  • Strong Linux/Unix administration skills
  • Good knowledge of Windows environments

Containerization & Cloud

  • Strong knowledge of Docker and Kubernetes
  • Understanding of cloud-based platforms and solutions

Infrastructure & Networking

  • Good understanding of enterprise infrastructure firewalls and networking concepts
  • Knowledge of load-balancing technologies
  • Strong understanding of networking fundamentals

Security

  • Experience with APIs
  • Familiarity with CyberArk or HashiCorp Vault

Databases

  • Experience with SQL Server
  • Experience with Oracle
  • Exposure to NoSQL databases

Monitoring & Observability

  • Experience configuring application monitoring tools such as Dynatrace

DevOps & CI/CD

Experience with:

  • Jenkins
  • Bitbucket
  • Artifactory
  • Ansible
  • ArgoCD

Development & Automation

  • Knowledge of software development and scripting methodologies
  • Demonstrated programming ability in languages such as Python

IT Service Management

  • Good understanding of ITIL processes
  • Understanding of user and server authentication mechanisms that enable automated deployment cycles while maintaining strong security controls

Personal Attributes

  • Strong problem-solving abilities
  • Team-oriented mindset
  • Customer-focused approach


Employment Type : Full Time
Experience: years
Vacancy: 1
Create a job alert for this search

Senior Site Reliability Engineer – Automation & Observability • Montreal, Quebec, Canada

Similar jobs

Site Reliability Engineer

Vertex Elite LLCRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Duration: ContractKey Skills:Monitoring / Observability tools - Dynatrace, ELK etc.Platform/ cloud Observability - OpenShift, Prometheus / Azure Cloud etc.Key Responsibilities:Collaborate with vari... Show more

 • Promoted

Radformation Senior Engineer - Adaptive Systems

Radformationmontreal (administrative region), qc, Canada
Full-time

Elevate your career as a Senior Engineer in Adaptive Systems at Radformation, working remotely to transform cancer care.Build software that enhances treatment efficiency for radiation oncology.The ... Show more

 • Promoted

Site Reliability Engineer

MaintainXMontréal, Canada
Full-time

MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 13,000+ customers including Duracell, Shell, Cintas, and Brenntag.We raised $150M in Series D fund... Show more

 • Promoted

Site Reliability Engineer

Hunter BondMontréal, Canada
Full-time

Role: DevOps EngineerClient: Most Elite Tech Firm in CanadaCompensation: Up to $200k CAD + Bonus + PackageLocation: MontrealOverviewAn Elite FinTech Firm is looking for a highly talented DevOps Eng... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsMontreal (administrative region), QC, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Intact Hybrid Site Reliability Engineer

IntactMontreal
Full-time

Join Intact as a Site Reliability Engineer and elevate operational reliability across cloud platforms.This hands-on role employs Azure, AWS, and GCP expertise to manage incidents effectively.The SR... Show more

 • Promoted

Senior Ii Site Reliability Engineer

Akamai TechnologiesRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Job DescriptionJoin our SRE team! Our team uses large datasets to analyze and measure the performance and reliability of our platform.We are networking data scientists: we combine our knowledge of ... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificMontreal (administrative region), QC, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Site Reliability Engineer

BasetenMontréal, Canada
Full-time

About BasetenBaseten powers mission‐critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting applied AI resear... Show more

 • Promoted

SAR Payload Systems Engineering Advisor

Sky Systems, Inc. (SkySys)montreal (administrative region), qc, Canada
Full-time

SAR Payload Systems Engineering Advisor.Duration: 12 months – 40 hours per week.Location: Montreal – West Island, 4 days per week in office.Salary: CAD $125K - $140K Annually with Standard Benefits... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Specialist in Site Reliability Engineering

Global Talent Alliance, CanadaMontreal
Full-time

Become a pivotal part of I&T solutions as a Senior Specialist Site Reliability Engineer.Focus on RAM analysis and reliability in complex cloud-based systems.With a primary emphasis on system robust... Show more

 • Promoted

Senior Engineering Developer, Site reliability (Hybrid)

National BankMontreal, Quebec
Full-time +2

Senior Engineering Developer, Site reliability.IT Delivery, Wealth Management sector at National Bank means acting as a specialist in application reliability, observability and operational maturity... Show more

Sr. Reliability Engineer - Products & Solutions - $93,000 - $133,000 A Year

Siemens HealthineersLe Plateau, Canada
Full-time

Seeking a Senior Reliability Engineer to model, test, and monitor product reliability, ensuring compliance with medical device regulations and supporting product development from design to manufact... Show more

 • Promoted

Site Reliability Engineer at mthree

mthree Recruiting PortalMontreal (administrative region), QC, CA
Full-time

Become a Site Reliability Engineer with mthree in Montreal, focusing on high-scale, reliable technology systems.Your role will emphasize system performance, availability, and scalability.As part of... Show more

 • Promoted

Senior Site Reliability Engineer

SecurityScorecardMontreal
Full-time

SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries.Founded in 2013 by security and risk experts Dr.Alex Ya... Show more

 • Promoted

Remote Site Reliability Engineer - Scale Crypto Systems

NewtonMontreal (administrative region), QC, CA
Remote
Full-time

A leading innovative tech company in Toronto is looking for a Site Reliability Engineer.In this pivotal role, you will enhance the reliability and resilience of critical services, manage incidents,... Show more

 • Promoted

Reliability Engineer - C$100,000 - C$150,000 A Year

Epiroc CanadaLe Plateau, Canada
Full-time

Reliability Engineer responsible for developing and implementing processes for product serviceability, reliability, and maintainability, performing data analytics, and communicating with stakeholders. Show more