Talent.com
Caseware
Remote Staff Site Reliability EngineerCaseware • Laval, Quebec
No longer accepting applications
Remote Staff Site Reliability Engineer

Remote Staff Site Reliability Engineer

Caseware • Laval, Quebec
15 days ago
Job type
  • Full-time
  • Permanent
  • Remote
Job description

Caseware is one of Canada's original Fintech companies, having led the global audit and accounting software industry for over 30 years, with more than 500,000 users across 130 countries and available in 16 different languages. While you might not have heard of us (yet) over 36,000 accounting and audit professionals list Caseware as a skill on their LinkedIn profiles!




This is a hands-on senior engineering role focused on improving production resilience, strengthening security, driving operational excellence, and enhancing the developer experience across the organization.


In this role, you will design, build, and evolve the foundational systems, tooling, and operational practices that enable engineering teams to ship secure, reliable, and scalable software with confidence. You will help establish reliability standards, define service level objectives (SLOs), improve observability, automate operational processes, and drive incident management and post-incident learning practices that strengthen platform stability over time.


Partnering closely with Engineering, Security, Platform, and Product teams, you will architect scalable distributed systems, optimize Kubernetes and AWS-based infrastructure, and build automated delivery pipelines that support rapid and safe software releases. You will play a key role in reducing operational toil, improving system performance, increasing platform reliability, and ensuring that our infrastructure can support continued business growth.






❗ This is a full-time permanent position


❗ This is an existing vacancy



Location: This is a remote location open to candidates legally authorized to work in Canada.



What you will be doing:



  • Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes.

  • Design, implement, and continuously improve deployment, release, and rollback strategies across complex distributed systems.

  • Establish secure-by-default CI/CD pipelines with robust automation, governance, and policy-driven controls.

  • Enhance platform observability through metrics, logs, tracing, and actionable alerting to improve system visibility and operational efficiency.

  • Define, implement, and mature Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability standards across the organization.

  • Lead response efforts for high-severity incidents, ensuring timely resolution, effective communication, and meaningful post-incident reviews that drive continuous improvement.

  • Partner closely with engineering teams to strengthen platform standards, improve service resilience, optimize runtime performance, and embed reliability best practices.

  • Mentor and guide engineers on cloud-native technologies, site reliability engineering principles, and operational excellence practices, fostering a culture of continuous learning and accountability.


What you will bring:


  • 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or related cloud-native engineering roles.

  • Deep expertise in AWS services, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security best practices.

  • Advanced experience operating and scaling production Kubernetes environments.

  • Strong hands-on experience with Istio service mesh, including traffic management, security, observability, and resiliency.

  • Proven expertise with Infrastructure as Code (IaC), preferably using AWS CDK.

  • Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms.

  • Strong troubleshooting, performance optimization, and incident management experience in distributed systems.

  • Excellent communication, collaboration, and technical leadership skills.


Observability & Reliability



  • Experience designing and operating monitoring, logging, tracing, and alerting solutions for cloud-native platforms.

  • Strong knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling.

  • Experience defining and operationalizing SLIs, SLOs, alerting strategies, runbooks, and reliability metrics.

  • Proven ability to leverage observability data to improve service reliability, reduce incident impact, and optimize operational performance.


Software Engineering & Platform Development



  • Strong proficiency in TypeScript and Node.js for platform engineering, automation, and operational tooling.

  • Experience building and maintaining scalable backend services, APIs, and event-driven systems.

  • Deep understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking.

  • Experience implementing zero-trust architectures, mTLS, and service-to-service security controls.

  • Commitment to high-quality engineering practices, including automated testing, code reviews, and observability-driven development.

  • Strong understanding of resilience engineering, including autoscaling, disruption management, failure testing, and safe deployment strategies.


Nice to Have



  • Experience with progressive delivery practices such as canary, blue/green, and feature-flag-based deployments.

  • Experience working in regulated, compliance-driven, or security-sensitive SaaS environments.

  • Familiarity with FinOps principles and cost optimization strategies for cloud platforms.

  • Experience building internal developer platforms and self-service engineering tooling.

  • Cloud-native certifications such as CKA, CKAD, CKS, KCSA, or KCNA.

  • Kubestronaut certification or equivalent advanced Kubernetes expertise is highly regarded.



Salary Range:


The annual base salary for this position is between $140,000 CAD and $155,000 CAD per year.


This role is also eligible for discretionary bonus and/or commission, as well as other benefits. Actual pay within the listed range will be determined based on factors such as transferable skills, relevant experience, market conditions, and primary work location. The posted range is subject to change and may be updated periodically.





What's in it for you:


▪️Innovation is at our core. We work with cutting-edge technology in accounting and financial reporting, constantly pushing the boundaries to create impactful software solutions.

▪️We are committed to a collaborative culture, where your ideas are valued, and knowledge sharing is encouraged within a supportive, inclusive team.

▪️Work-life balance is important to us. We offer flexible work options, remote opportunities, and generous time-off policies to ensure a healthy work-life balance.

▪️We offer competitive compensation, including a competitive salary and comprehensive benefits such as health insurance and retirement plans.

▪️We are driven by impactful work. Your contributions directly affect how our clients manage financial processes and drive their success.

▪️Recognition and rewards matter to us. We celebrate hard work through recognition programs, performance bonuses, and opportunities for career growth.

▪️We embrace global opportunities. Work on international projects and collaborate with a diverse, global team.


About Caseware:

Caseware's cutting-edge software products are meticulously designed for accounting firms, corporations, and governments. Our teams are continually collaborating, innovating, and building upon our existing suite of products. With a customer-focused mindset, we are building technology that is shaping what the future of audits, financial reporting, and financial data analytics will look like.


With a recent strategic investment from Hg Capital in 2020, Caseware is now in its next major growth phase as we double down on the people and products that have made Caseware so successful to date.


One of Caseware's core values is Many Voices, One Team and with that in mind, we're dedicated to building teams as diverse as our customers in an equitable and inclusive way. We welcome and encourage candidates of all backgrounds to apply. Should you require accommodations or have any questions at any point during the application or interview process, please e-mail our People Operations team at [email protected].


AI Usage:

The recruitment process may use AI assisted tools but not for candidate screening or assessment. All final hiring decisions are made by humans to ensure fairness, transparency, and oversight.


Background Check:

Any candidates successful in obtaining an offer for a position will need to successfully complete a background check through Certn.co which typically includes an Identity Verification and Criminal Record Check. Executives and Senior Managers will undergo a Soft Credit Check as well. Candidates residing in the Netherlands and Germany are excluded from undergoing background checks via Certn.co


Security and Fraud:

Caseware takes the security of candidates seriously. All legitimate communication from us will come from email addresses ending in @caseware.com and our open positions are always listed on reputable job boards and on our website https://jobs.lever.co/caseware. We will NEVER ask for payment or financial information from you. If you receive an unsolicited job offer, proceed with extreme caution.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Create a job alert for this search

Remote Staff Site Reliability Engineer • Laval, Quebec

Similar jobs

Site Reliability Engineer

MaintainXMontreal (administrative region), QC, CA
Full-time

MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 13,000+ customers including Duracell, Shell, Cintas, and Brenntag.We raised $150M in Series D funding ... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificMontreal (administrative region), QC, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Senior Full-Stack Engineer - Accessibility & Inclusive Tech (Remote)

Accessibility Partners CanadaMontreal (administrative region), QC, CA
Remote
Full-time

A leader in accessible technology is seeking a Senior Full-Stack Developer to create equitable and accessible digital systems.This role involves developing both front-end and back-end systems, focu... Show more

 • Promoted

Experienced Site Reliability Engineer Remote

Tecsys Inc.Montreal (administrative region), QC, CA
Remote
Full-time

Join Tecsys as an Experienced Site Reliability Engineer and elevate our cloud infrastructure reliability.Work remotely and focus on automation and system health.At Tecsys, we are searching for a Si... Show more

 • Promoted

Site Reliability Engineer

Tecsys Inc.Montreal (administrative region), QC, CA
Permanent

Having recognized the advantages of remote work, including employee morale, productivity, reduced commuting on employee wellbeing and the environment, we are proud to be a digital-first company.The... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Site Reliability Engineer (SRE), ServiceNow, Application Infrastructure

Open Systems TechnologiesMontreal, Montreal (administrative region), CA
Full-time

Site Reliability Engineer (SRE), ServiceNow, Application Infrastructure.Be among the first 25 applicants.The Application Infrastructure (AI) department is seeking a Site Reliability Engineer (SRE) ... Show more

 • Promoted

Site Reliability Engineer - C$125,000 - C$250,000 A Year

ApTaskMont-Royal, Canada
Full-time

The Site Reliability Engineer will focus on improving system reliability and performance within a ServiceNow SaaS implementation, including troubleshooting and automation. Show more

 • Promoted

Senior Tailings Engineer - Remote Leadership & Design

StantecMontreal (administrative region), QC, CA
Remote
Full-time

A leading engineering firm is seeking an experienced engineer specializing in mine tailings management to join their Montreal team.This pivotal role involves leading complex technical projects, dev... Show more

 • Promoted

Site Reliability Engineer for Cloud Infrastructure Management

NewtonMontreal (administrative region), QC, CA
Full-time

Be a pivotal Site Reliability Engineer focused on improving infrastructure resilience and reliability.Collaborate remotely to drive operational success and enhance system performance in a dynamic e... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificMontreal (administrative region), QC, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT GroupMontreal, Montreal (administrative region), CA
Full-time

Site Reliability Engineer (Linux / Cloud Infrastructure) role with hands-on experience across Linux, distributed systems, scripting, databases, monitoring, containers, cloud SaaS integrations, mess... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Site Reliability Engineer

BasetenMontreal (administrative region), QC, CA
Full-time

Baseten powers mission‑critical inference for the world’s most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting applied AI research, flexible infr... Show more

 • Promoted

DevOps/Site Reliability Engineer - Up to $200k CAD + Bonus - Elite Tech Firm

Hunter BondMontreal (administrative region), QC, CA
Full-time

DevOps/Site Reliability Engineer.Most Elite Tech Firm in Canada.Up to $200k CAD + Bonus + Full Package.One of Canada’s most elite tech firms is hiring a Site Reliability Engineer to join a seriousl... Show more

 • Promoted

Senior Site Reliability Engineer (SRE) – Automation & Observability

Tech Talent InternationalMontreal (administrative region), QC, CA
Permanent

Senior Site Reliability Engineer (SRE) – Automation & Observability.Job Openings Senior Site Reliability Engineer (SRE) – Automation & Observability.About the job Senior Site Reliability Engineer (... Show more

 • Promoted

Site Reliability Engineer

ApTaskMontreal, Montreal (administrative region), CA
Full-time

Direct message the job poster from ApTask.Looking for an intermediate between 2 to 5 years' experience.The Application Infrastructure (Al) department is seeking a Site Reliability Engineer (SRE) to... Show more

 • Promoted

Remote Site Reliability Engineer - Scale Crypto Systems

NewtonMontreal (administrative region), QC, CA
Remote
Full-time

A leading innovative tech company in Toronto is looking for a Site Reliability Engineer.In this pivotal role, you will enhance the reliability and resilience of critical services, manage incidents,... Show more

 • Promoted

Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc.Montreal (administrative region), QC, CA
Full-time

A leading supply chain solutions provider is seeking a Site Reliability Engineer to optimize and ensure the reliability of their cloud infrastructure across AWS and Kubernetes.This role emphasizes ... Show more