Talent.com
Lightspeed
Staff Site Reliability EngineerLightspeed • Montreal
Staff Site Reliability Engineer

Staff Site Reliability Engineer

Lightspeed • Montreal
14 days ago
Job type
  • Full-time
Job description

Hi there! Thanks for stopping by 👋

Are you actively looking for a new opportunity? Or just checking the market? Well… you might just be in the right place!

We’re looking for a Staff Site Reliability Engineer to join our Data team in Canada.

As a Staff Data SRE, you are the technical backbone of the Data Office's infrastructure platform. Your scope spans the entire Data Business Unit — you solve systemic, platform-level problems, not individual tickets. You own the reliability, scalability, and developer experience of the data platform, and you act as a force multiplier: making the engineers around you faster, the systems more resilient, and the platform easier to consume.

You bring deep GCP expertise and a product mindset to data infrastructure. You are comfortable navigating ML/AI workload requirements (Vertex AI, feature stores, training pipelines) and can reason confidently about the BI tooling and data layers that feed it (Looker, BigQuery). Your decisions carry weight across teams, and you communicate them clearly to both engineers and non-technical stakeholders.

Please note this role is based in Montreal, Canada

What you’ll be doing:

  • You will own and be accountable for advancing the Data Office infrastructure engineering practices and delivering high-leverage platform projects. Your time will be distributed as follows:

    70% – Hands-On Engineering & Platform Ownership: Design and deliver significant infrastructure improvements. You write production-quality IaC (Terraform), lead solutions from design through delivery, and are the person the team calls for the hardest platform problems. This includes data infrastructure for batch/streaming workloads, ML/AI environments (Vertex AI, model serving, GPU-backed compute), and the BI serving layer (Looker infrastructure and GCP integration).

    10% – Team Coordination & Technical Leadership: Lead architecture and design discussions for larger cross-team projects. You review critical PRs, drive solution design sessions, and provide technical direction that keeps the team coherent and moving forward.

    10% – General Meetings: Lead conversations in agile ceremonies, incident reviews, and cross-functional syncs. You shift from participant to driver.

    10% – Mentorship & Technical Review: Actively mentor Senior and Intermediate SREs. You set the bar for IaC quality, observability practices, and operational discipline through reviews, pairing, and documentation.

  • Avoiding Pitfalls: You proactively identify risks in complex data migrations or infrastructure changes and propose mitigation strategies to ensure zero data loss and minimal downtime.

And a little bit of…

  • Participating in on-call rotation and incident response.

  • Contributing as part of the wider team to achieve organization-wide objectives even if this means doing things that aren't strictly within the scope of your role.

What you need to bring:

  • GCP Infrastructure: Deep expertise across GCP compute, networking, IAM, GKE, data services, and FinOps.

  • IaC Proficiency: Terraform as primary tool.

  • Hands-on experience with Looker infrastructure and ML/AI platform tooling (Vertex AI, model serving, training pipelines).

  • Proficient in Bash and Golang; Python a plus for data tooling.

  • Strong experience with metrics, logs, traces, alerting, and SLO/SLI design, DataDog,...

  • Product Mindset: You focus on making the platform easy to use, not just "available."

  • Experience with Github actions, Circle CI, GCP Cloud Build, …

  • Ability to communicate infrastructure trade-offs to both engineers and non-technical stakeholders.

  • AI proficiency, Go-to AI expertise within the Data BU — evaluates tooling, drives cross-team AI-first development practices.

  • Security-First Mindset, every design decision evaluated through a security and compliance lens.

  • Self-awareness with a willingness to learn and improve.

  • Ability to mentor and train other Engineers in the team.

We know that people are more than what’s on their CV. If you’re unsure that you have the right profile for the role... hit the ‘Apply’ button and give it a try!

Be a changemaker

You’ll enjoy:

  • A flexible work environment that empowers you to do your best work

  • A culture that celebrates performance

  • The chance to make an impact in a team that’s big enough for career growth, but lean enough to make your voice heard

  • Career-defining opportunities

Plus benefits designed to keep you happy, healthy and fulfilled.

  • Flexible paid time off and remote work policies

  • Equity options, because this is your company too

  • Contributions to your pension plan. Your future matters

  • Training opportunities to grow your skills and career

  • Health and wellness credit so you feel your best

  • Time off to volunteer and give back to your community

  • Interest groups, employee led networks, social committees to sponsored sports teams

  • Computer purchase program to get your personal Macbook

  • Enhanced parental leave to support growing families

Fuel your growth. Find your people.

At Lightspeed, your growth is our priority. We invest in you with continuous learning opportunities, global mobility and benefits designed to support you—all within a driven, diverse and inclusive team that’s passionate about empowering our communities

Please note that we ask applicants to disclose any criminal convictions, and we conduct criminal record checks as part of our hiring process for this role.

Create a job alert for this search

Staff Site Reliability Engineer • Montreal

Similar jobs

Site Reliability Engineer

ApTaskMontréal, Quebec, Canada
Full-time

Direct message the job poster from ApTask Looking for an intermediate between 2 to 5 years' experience.The Application Infrastructure (Al) department is seeking a Site Reliability Engineer (SRE) to... Show more

 • Promoted

Site Reliability Engineer

Vertex Elite LLCRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Duration: ContractKey Skills:Monitoring / Observability tools - Dynatrace, ELK etc.Platform/ cloud Observability - OpenShift, Prometheus / Azure Cloud etc.Key Responsibilities:Collaborate with vari... Show more

 • Promoted

Sr. Engineer

TechDoQuestmontreal (administrative region), qc, Canada
Full-time

Perform icing numerical simulations on complex aerodynamic configurations.Prepare, execute, and analyze high‑lift and icing wind tunnel test campaigns, including CFD and certification.Architect and... Show more

 • Promoted

Senior Full-Stack Engineer - Accessibility & Inclusive Tech (Remote)

Accessibility Partners CanadaMontreal (administrative region), QC, CA
Remote
Full-time

A leader in accessible technology is seeking a Senior Full-Stack Developer to create equitable and accessible digital systems.This role involves developing both front-end and back-end systems, focu... Show more

 • Promoted

Site Reliability Engineer

Hunter BondMontréal, Canada
Full-time

Role: DevOps EngineerClient: Most Elite Tech Firm in CanadaCompensation: Up to $200k CAD + Bonus + PackageLocation: MontrealOverviewAn Elite FinTech Firm is looking for a highly talented DevOps Eng... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsMontreal (administrative region), QC, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Senior Ii Site Reliability Engineer

Akamai TechnologiesRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Job DescriptionJoin our SRE team! Our team uses large datasets to analyze and measure the performance and reliability of our platform.We are networking data scientists: we combine our knowledge of ... Show more

 • Promoted

Senior Tailings Engineer - Remote Leadership & Design

StantecMontreal (administrative region), QC, CA
Remote
Full-time

A leading engineering firm is seeking an experienced engineer specializing in mine tailings management to join their Montreal team.This pivotal role involves leading complex technical projects, dev... Show more

 • Promoted

Remote Full-Stack Engineer for SaaS Platform | High Impact

Enso Connect Inc.Montreal (administrative region), QC, CA
Remote
Full-time

A tech startup in hospitality seeks a full stack software developer to join its growing team.This remote role involves significant ownership over work and the chance to impact the guest experience ... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificMontreal (administrative region), QC, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Site Reliability Engineer

BasetenMontréal, Canada
Full-time

About Baseten Baseten powers mission‑critical inference for the world’s most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting applied AI research,... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Engineering Sr. Advisor_1ENGEK

TRACTEBELmontreal (administrative region), qc, Canada
Full-time

TRACTEBEL, ranked 1st in Hydropower design & 4th in Nuclear design by Engineering News-Record’s (ENR) 2025 annual ranking, is a global engineering and consulting company dedicated to engineering a ... Show more

 • Promoted

Senior Infrastructure Reliability Engineer

ShippoMontreal (administrative region), QC, CA
Full-time

Enhance shipping solutions as a Senior Site Reliability Engineer in a remote setting.Focus on infrastructure integrity, scalability, and performance in a collaborative environment.This position inv... Show more

 • Promoted

Senior Site Reliability Engineer

SecurityScorecardmontreal (administrative region), qc, Canada
Full-time

SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries.Founded in 2013 by security and risk experts Dr.Alex Ya... Show more

 • Promoted

Site Reliability Engineer

MaintainXMontreal
Full-time

MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 13,000+ customers including Duracell, Shell, Cintas, and Brenntag.We raised $150M in Series D funding ... Show more

 • Promoted

Site Reliability Engineer - Tech Talent International

Tech Talent InternationalMontreal
Full-time

Join Tech Talent International as a Senior Site Reliability Engineer, specializing in Automation & Observability, located in Montreal.This hybrid role focuses on enhancing production efficiency and... Show more

 • Promoted

Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT GroupMontréal, Quebec, Canada
Full-time

Site Reliability Engineer (Linux / Cloud Infrastructure) role with hands-on experience across Linux, distributed systems, scripting, databases, monitoring, containers, cloud SaaS integrations, mess... Show more