Talent.com
Parallel Domain
Remote Senior Site Reliability EngineerParallel Domain • Georgetown, Ontario
Remote Senior Site Reliability Engineer

Remote Senior Site Reliability Engineer

Parallel Domain • Georgetown, Ontario
30+ days ago
Job type
  • Full-time
  • Remote
Job description

About the Role


Before an autonomous vehicle navigates a busy intersection, before a robot learns to pick and place in a warehouse, before any Physical AI system is trusted in the real world, it has to prove itself in ours. Parallel Domain builds the platform that validates the next generation of autonomous systems in high-fidelity virtual environments, and the infrastructure underneath that platform is what makes simulation at scale possible.


We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits at the core of how we run large-scale, distributed simulation workloads for autonomous-systems testing and validation. You'll work across multi-region AWS infrastructure, operate Kubernetes at scale, and contribute directly to reliability, security, and deployment systems that the rest of the engineering org depends on.


This is a hands-on role with the broad ownership typical of a startup. You'll partner closely with platform, simulation, and ML teams to keep the system running smoothly and evolving. We're growing the team—two of these roles are open—and the work is substantive: multi-region GPU scheduling, Windows workloads on Kubernetes, large-scale batch simulation, and an enterprise product direction that will require rethinking parts of how we deploy and operate.


Responsibilities



  • Infrastructure ownership and cloud operations. Design, build, and maintain multi-region AWS infrastructure using Terraform. Operate and scale EKS clusters across production regions: autoscaling, node lifecycle, workload health. Manage networking across environments: VPC design, DNS, load balancing, and cross-region connectivity. Support infrastructure changes, migrations, and expansions into new regions. Contribute to and improve GitOps-based deployment workflows using GitHub Actions, Helm, and Kustomize.




  • Reliability engineering and incident response. Help build and run incident management processes: severity definitions, escalation paths, on-call practices. Lead incident response, debugging, and root-cause analysis. Write postmortems and drive systemic reliability improvements from what they surface. Improve observability across metrics, logging, tracing, and dashboards. Support GPU and batch workloads running on Kubernetes.




  • Security and access management. Provide security-conscious feedback on platform architecture decisions. Own cloud IAM governance: roles, policies, and access boundaries across accounts and services. Lead compliance-adjacent work including audit-readiness, partner certification requirements, and supporting responses to customer security questionnaires.



  • Platform tooling and developer experience. Improve CI/CD pipelines and infrastructure validation. Support engineers with infrastructure debugging, environment setup, and performance issues. Contribute to tooling and automation in Python and Bash. Take on adjacent responsibilities as needed in a startup environment.


Required Qualifications



  • Experience. 5+ years in SRE, DevOps, or infrastructure engineering roles, with a track record of operating production systems across multiple regions.




  • Terraform. Modules, state management, and multi-environment patterns.




  • AWS depth. Solid experience across VPC, IAM, EKS, S3, and CloudWatch.




  • Kubernetes expertise. Cluster operations, autoscaling, RBAC, and Helm.




  • CI/CD and GitOps. Experience with GitHub Actions, ArgoCD, or similar workflows.




  • Networking fundamentals. CIDR, DNS, load balancing, VPN, and cross-region connectivity.




  • Observability. Experience with tooling such as Prometheus and Grafana.




  • Scripting. Comfort with Python and Bash for tooling and automation.




  • Cross-platform familiarity. Working knowledge of both Linux and Windows environments. Operational experience supporting Windows-based workloads is a meaningful advantage.



  • Pragmatism and ownership. Comfortable in a fast-moving startup with evolving priorities. You take ownership of systems while collaborating closely with other teams, and you're pragmatic about tradeoffs between speed, reliability, and complexity.


Preferred Qualifications



  • Windows on Kubernetes. Experience with Windows node pools, Windows AMIs, and GPU-adjacent components on K8s.




  • GPU scheduling. Familiarity with GPU scheduling on Kubernetes, including NVIDIA device plugin configuration.




  • Domain workloads. Experience supporting simulation, ML, or rendering workloads in cloud infrastructure.




  • AWS extras. Exposure to AWS Storage Gateway, Active Directory integrations, or AWS Transfer Family.




  • Service mesh. Familiarity with service proxy or service mesh patterns.




  • Container OS. Experience with container-optimized OS images (e.g., Bottlerocket, Packer).



  • Cost optimization. Cloud cost optimization at scale.


Core Tools

Terraform · AWS · Kubernetes · Helm · Kustomize · ArgoCD · GitHub Actions · Prometheus · Grafana · Docker · Python · Bash

What Makes a Great Candidate

You think in failure modes and proactively surface issues. You hold a principled view on security and push back constructively when designs introduce unnecessary risk. You communicate clearly across engineering, product, and customer-facing teams, flagging issues with urgency proportional to customer impact. You take end-to-end ownership of complex efforts and know when to push for the clean solution versus the pragmatic one.Base salary range of CAD $145,000–$185,000, depending on skills, qualifications, and experience, plus equity, full health/dental/vision coverage, learning stipend, and generous vacation. This role is remote-friendly across Canada and the US Pacific Northwest.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Create a job alert for this search

Remote Senior Site Reliability Engineer • Georgetown, Ontario

Similar jobs

Senior RAMS Engineer | Rail Safety & Reliability Expert

Egis in CanadaHamilton, ON, CA
Full-time

A leading engineering firm in Hamilton, Ontario, is seeking a qualified RAMS Engineer to support the development and execution of the RAMS program.The ideal candidate will have a Bachelor's degree ... Show more

 • Promoted

Entry-Level Mechanical Engineering Role in Innovative Designs

L3Harris TechnologiesHamilton, ON, CA
Full-time

Launch your career as an Entry-Level Mechanical Engineer, creating sophisticated designs for defense applications.Contribute to impactful projects while enjoying a unique work schedule.In this posi... Show more

 • Promoted

Applied Systems Engineer

L3Harris TechnologiesHamilton, ON, CA
Full-time

Applied Systems Engineer Specialist (L3).Employees work 9 out of every 14 days – totaling 80 hours worked – and have every other Friday off.Acts as the primary technical interface between Wescam pr... Show more

 • Promoted

Remote Staff Engineer: Capacity Modeling Expertise

Affirmhamilton, on, Canada
Remote
Full-time

Become part of Affirm's innovative approach as a Staff Software Engineer in Capacity Modeling, focused on translating business demand into technical capacity solutions.This remote role highlights y... Show more

 • Promoted

Remote Containerization & Virtualisation Engineer

Canonicalhamilton, on, Canada
Remote
Full-time

A leading provider of open source technology in Canada seeks exceptional software engineers with expertise in Go, Rust, or C/C++.You’ll contribute to advancement in virtualization and container tec... Show more

 • Promoted

Remote CIAM Innovation Engineer

Affirmhamilton, on, Canada
Remote
Full-time

Unlock your potential as a Remote CIAM Innovation Engineer, specializing in customer identity and authentication solutions.Your engineering prowess will enhance backend services that enrich account... Show more

 • Promoted

Innovative Engineer for Next-Gen Containerization and Virtualization

Canonicalhamilton, on, Canada
Full-time

Lead the charge in containerization and virtualization technology as a talented engineer.Engage in exciting projects with cutting-edge technologies while enjoying a fully remote work environment.Th... Show more

 • Promoted

Senior Systems Engineer, Defense Integration - $85,500 - $135,500 A Year

WescamHamilton, Canada
Full-time

A leading defense technology firm in Ontario is seeking an Applied Systems Engineer Specialist to act as the primary technical interface between Wescam products and customers.The role involves mana... Show more

 • Promoted

Reliability Engineer Hamilton, ON 7/6/2026

Maple Leaf FoodsHamilton, ON, CA
Full-time

Reporting to the Reliability Manager, this position will be responsible for developing and improving manufacturing equipment and processes by the application of engineering principles with a LEAN a... Show more

 • Promoted

Containerization & Virtualisation Engineer - Remote

CanonicalHamilton, Canada
Remote
Full-time

A leading provider of open source technology in Canada seeks exceptional software engineers with expertise in Go, Rust, or C/C++.You'll contribute to advancement in virtualization and container... Show more

 • Promoted

Reliability Engineer

Maple Leaf Foods IncHamilton, ON, CA
Full-time

Reporting to the Reliability Manager, this position will develop and improve manufacturing equipment and processes by applying engineering principles with a LEAN approach.Compensation: $71,000 - $1... Show more

 • Promoted

Senior Waste Management Engineer in Ontario

Dillon Consulting LimitedHamilton, ON, CA
Full-time

Advance your career as a Senior Waste Management Engineer with Dillon in Ontario.Drive project innovation and engage with clients while showcasing your expertise.As a Senior Waste Management Engine... Show more

 • Promoted

Senior Mining Permitting Lead (Remote)

Stantec Consulting International Ltd.Hamilton, ON, CA
Remote
Full-time

A key environmental consulting firm is seeking a Senior Mining Practitioner to join their environmental permitting practice.This role involves leading environmental permitting, regulatory complianc... Show more

 • Promoted

Reliability Engineering Manager

itec group Inc.hamilton, on, Canada
Full-time

We’re excited to share an opportunity for an existing vacancy as a Manager of Reliability.This role is responsible for leading plant reliability engineering, preventative maintenance, and associate... Show more

 • Promoted

Engineering Role: Reliability in Manufacturing

Maple Leaf Foods IncHamilton, ON, CA
Full-time

Become a key player in manufacturing reliability as a Reliability Engineer applying LEAN methodologies.Help improve processes and enhance efficiency across the plant.You will work under the Reliabi... Show more

 • Promoted

Senior Software Engineering Manager Remote

AffirmHamilton, ON, CA
Remote
Full-time

As a Senior Engineering Manager at Affirm, focus on amplifying developer productivity and innovation across teams.This remote role emphasizes leadership in environment management and system improve... Show more

 • Promoted

Engineering Role in Reliability Enhancement

Maple Leaf FoodsHamilton, ON, CA
Full-time

Drive reliability and safety as a Reliability Engineer at Maple Leaf Foods, utilizing engineering principles to enhance manufacturing equipment and processes through a LEAN approach.This engineerin... Show more

 • Promoted

Software Engineer II, Backend (Reliability Platform)

United States Digital Space LLCHamilton, ON, CA
Full-time

Chief of Staff to the Chief Product Officer.As Chief of Staff to the Chief Product Officer, you will be a force multiplier for the company's product leadership team.This role helps the CPO run a hi... Show more

 • Promoted

Remote Engineering Manager for MAAS

CanonicalHamilton, ON, CA
Remote
Full-time

Serve as the Engineering Manager for MAAS at Canonical, supporting a fully remote team throughout EMEA or the Americas.Guide the growth of private bare-metal infrastructure while mentoring your eng... Show more

 • Promoted

Backend Engineer, Ai Agents - Remote - $125,000 - $175,000 A Year - Remote

Financial technology companyHamilton, Canada
Remote
Full-time

Seeking a Software Engineer II for an AI Agents team in a Canadian FinTech company to develop APIs and contribute to product development. Show more