Talent.com
Parallel Domain
Remote Senior Site Reliability EngineerParallel Domain • Chateauguay, Quebec
Remote Senior Site Reliability Engineer

Remote Senior Site Reliability Engineer

Parallel Domain • Chateauguay, Quebec
30+ days ago
Job type
  • Full-time
  • Remote
Job description

About the Role


Before an autonomous vehicle navigates a busy intersection, before a robot learns to pick and place in a warehouse, before any Physical AI system is trusted in the real world, it has to prove itself in ours. Parallel Domain builds the platform that validates the next generation of autonomous systems in high-fidelity virtual environments, and the infrastructure underneath that platform is what makes simulation at scale possible.


We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits at the core of how we run large-scale, distributed simulation workloads for autonomous-systems testing and validation. You'll work across multi-region AWS infrastructure, operate Kubernetes at scale, and contribute directly to reliability, security, and deployment systems that the rest of the engineering org depends on.


This is a hands-on role with the broad ownership typical of a startup. You'll partner closely with platform, simulation, and ML teams to keep the system running smoothly and evolving. We're growing the team—two of these roles are open—and the work is substantive: multi-region GPU scheduling, Windows workloads on Kubernetes, large-scale batch simulation, and an enterprise product direction that will require rethinking parts of how we deploy and operate.


Responsibilities



  • Infrastructure ownership and cloud operations. Design, build, and maintain multi-region AWS infrastructure using Terraform. Operate and scale EKS clusters across production regions: autoscaling, node lifecycle, workload health. Manage networking across environments: VPC design, DNS, load balancing, and cross-region connectivity. Support infrastructure changes, migrations, and expansions into new regions. Contribute to and improve GitOps-based deployment workflows using GitHub Actions, Helm, and Kustomize.




  • Reliability engineering and incident response. Help build and run incident management processes: severity definitions, escalation paths, on-call practices. Lead incident response, debugging, and root-cause analysis. Write postmortems and drive systemic reliability improvements from what they surface. Improve observability across metrics, logging, tracing, and dashboards. Support GPU and batch workloads running on Kubernetes.




  • Security and access management. Provide security-conscious feedback on platform architecture decisions. Own cloud IAM governance: roles, policies, and access boundaries across accounts and services. Lead compliance-adjacent work including audit-readiness, partner certification requirements, and supporting responses to customer security questionnaires.



  • Platform tooling and developer experience. Improve CI/CD pipelines and infrastructure validation. Support engineers with infrastructure debugging, environment setup, and performance issues. Contribute to tooling and automation in Python and Bash. Take on adjacent responsibilities as needed in a startup environment.


Required Qualifications



  • Experience. 5+ years in SRE, DevOps, or infrastructure engineering roles, with a track record of operating production systems across multiple regions.




  • Terraform. Modules, state management, and multi-environment patterns.




  • AWS depth. Solid experience across VPC, IAM, EKS, S3, and CloudWatch.




  • Kubernetes expertise. Cluster operations, autoscaling, RBAC, and Helm.




  • CI/CD and GitOps. Experience with GitHub Actions, ArgoCD, or similar workflows.




  • Networking fundamentals. CIDR, DNS, load balancing, VPN, and cross-region connectivity.




  • Observability. Experience with tooling such as Prometheus and Grafana.




  • Scripting. Comfort with Python and Bash for tooling and automation.




  • Cross-platform familiarity. Working knowledge of both Linux and Windows environments. Operational experience supporting Windows-based workloads is a meaningful advantage.



  • Pragmatism and ownership. Comfortable in a fast-moving startup with evolving priorities. You take ownership of systems while collaborating closely with other teams, and you're pragmatic about tradeoffs between speed, reliability, and complexity.


Preferred Qualifications



  • Windows on Kubernetes. Experience with Windows node pools, Windows AMIs, and GPU-adjacent components on K8s.




  • GPU scheduling. Familiarity with GPU scheduling on Kubernetes, including NVIDIA device plugin configuration.




  • Domain workloads. Experience supporting simulation, ML, or rendering workloads in cloud infrastructure.




  • AWS extras. Exposure to AWS Storage Gateway, Active Directory integrations, or AWS Transfer Family.




  • Service mesh. Familiarity with service proxy or service mesh patterns.




  • Container OS. Experience with container-optimized OS images (e.g., Bottlerocket, Packer).



  • Cost optimization. Cloud cost optimization at scale.


Core Tools

Terraform · AWS · Kubernetes · Helm · Kustomize · ArgoCD · GitHub Actions · Prometheus · Grafana · Docker · Python · Bash

What Makes a Great Candidate

You think in failure modes and proactively surface issues. You hold a principled view on security and push back constructively when designs introduce unnecessary risk. You communicate clearly across engineering, product, and customer-facing teams, flagging issues with urgency proportional to customer impact. You take end-to-end ownership of complex efforts and know when to push for the clean solution versus the pragmatic one.Base salary range of CAD $145,000–$185,000, depending on skills, qualifications, and experience, plus equity, full health/dental/vision coverage, learning stipend, and generous vacation. This role is remote-friendly across Canada and the US Pacific Northwest.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Create a job alert for this search

Remote Senior Site Reliability Engineer • Chateauguay, Quebec

Similar jobs

Site Reliability Engineer

Basetenmontreal (administrative region), qc, Canada
Full-time

Baseten powers mission‑critical inference for the world’s most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting applied AI research, flexible infr... Show more

 • Promoted

DevOps/Site Reliability Engineer - Up to $200k CAD + Bonus - Elite Tech Firm

Hunter Bondmontreal (administrative region), qc, Canada
Full-time

DevOps/Site Reliability Engineer.Most Elite Tech Firm in Canada.Up to $200k CAD + Bonus + Full Package.One of Canada’s most elite tech firms is hiring a Site Reliability Engineer to join a seriousl... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificMontreal (administrative region), QC, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Site Reliability Engineer

MaintainXMontreal (administrative region), QC, CA
Full-time

MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 13,000+ customers including Duracell, Shell, Cintas, and Brenntag.We raised $150M in Series D funding ... Show more

 • Promoted

Experienced Site Reliability Engineer Remote

Tecsys Inc.Montreal (administrative region), QC, CA
Remote
Full-time

Join Tecsys as an Experienced Site Reliability Engineer and elevate our cloud infrastructure reliability.Work remotely and focus on automation and system health.At Tecsys, we are searching for a Si... Show more

 • Promoted

Site Reliability Engineer

Tecsys Inc.Montreal (administrative region), QC, CA
Permanent

Having recognized the advantages of remote work, including employee morale, productivity, reduced commuting on employee wellbeing and the environment, we are proud to be a digital-first company.The... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsMontreal (administrative region), QC, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Senior Specialist in Site Reliability Engineering

Global Talent Alliance, CanadaMontreal (administrative region), QC, CA
Full-time

Become a pivotal part of I&T solutions as a Senior Specialist Site Reliability Engineer.Focus on RAM analysis and reliability in complex cloud-based systems.With a primary emphasis on system robust... Show more

 • Promoted

Site Reliability Engineer

Hunter Bondmontreal (administrative region), qc, Canada
Full-time

Most Elite Tech Firm in Canada.Up to $200k CAD + Bonus + Package.An Elite FinTech Firm is looking for a highly talented DevOps Engineer/Systems SRE to join a talented flat-structured team within a ... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificMontreal (administrative region), QC, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Specialist Site Reliability Engineer

Global Talent Alliance, CanadaMontreal (administrative region), QC, CA
Full-time

About the job Specialist Site Reliability Engineer.The role of the Specialist Site Reliability Engineer (SRE) is to execute RAM analysis and engineering in support of the I&T solutions.The overall ... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Engineering Developer, Site reliability

National Bank of Canadamontreal (administrative region), qc, Canada
Full-time

A career as a Senior Engineering Developer, Site reliability within the IT Delivery, Wealth Management sector at National Bank means acting as a specialist in application reliability, observability... Show more

 • Promoted

Senior Site Reliability Engineer

SecurityScorecardmontreal (administrative region), qc, Canada
Full-time

SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries.Founded in 2013 by security and risk experts Dr.Alex Ya... Show more

 • Promoted

Intact Hybrid Site Reliability Engineer

IntactMontreal (administrative region), QC, CA
Full-time

Join Intact as a Site Reliability Engineer and elevate operational reliability across cloud platforms.This hands-on role employs Azure, AWS, and GCP expertise to manage incidents effectively.The SR... Show more

 • Promoted

Site Reliability Engineer

ApTaskmontreal, montreal (administrative region), Canada
Full-time

Direct message the job poster from ApTask.Looking for an intermediate between 2 to 5 years' experience.The Application Infrastructure (Al) department is seeking a Site Reliability Engineer (SRE) to... Show more

 • Promoted

Remote Site Reliability Engineer - Scale Crypto Systems

NewtonMontreal (administrative region), QC, CA
Remote
Full-time

A leading innovative tech company in Toronto is looking for a Site Reliability Engineer.In this pivotal role, you will enhance the reliability and resilience of critical services, manage incidents,... Show more

 • Promoted

Site Reliability Engineer - Tech Talent International

Tech Talent InternationalMontreal (administrative region), QC, CA
Full-time

Join Tech Talent International as a Senior Site Reliability Engineer, specializing in Automation & Observability, located in Montreal.This hybrid role focuses on enhancing production efficiency and... Show more