Talent.com
Parallel Domain
Remote Senior Site Reliability EngineerParallel Domain • Markham, Ontario
Remote Senior Site Reliability Engineer

Remote Senior Site Reliability Engineer

Parallel Domain • Markham, Ontario
Il y a plus de 30 jours
Type de contrat
  • Temps plein
  • Télétravail
Description de poste

About the Role


Before an autonomous vehicle navigates a busy intersection, before a robot learns to pick and place in a warehouse, before any Physical AI system is trusted in the real world, it has to prove itself in ours. Parallel Domain builds the platform that validates the next generation of autonomous systems in high-fidelity virtual environments, and the infrastructure underneath that platform is what makes simulation at scale possible.


We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits at the core of how we run large-scale, distributed simulation workloads for autonomous-systems testing and validation. You'll work across multi-region AWS infrastructure, operate Kubernetes at scale, and contribute directly to reliability, security, and deployment systems that the rest of the engineering org depends on.


This is a hands-on role with the broad ownership typical of a startup. You'll partner closely with platform, simulation, and ML teams to keep the system running smoothly and evolving. We're growing the team—two of these roles are open—and the work is substantive: multi-region GPU scheduling, Windows workloads on Kubernetes, large-scale batch simulation, and an enterprise product direction that will require rethinking parts of how we deploy and operate.


Responsibilities



  • Infrastructure ownership and cloud operations. Design, build, and maintain multi-region AWS infrastructure using Terraform. Operate and scale EKS clusters across production regions: autoscaling, node lifecycle, workload health. Manage networking across environments: VPC design, DNS, load balancing, and cross-region connectivity. Support infrastructure changes, migrations, and expansions into new regions. Contribute to and improve GitOps-based deployment workflows using GitHub Actions, Helm, and Kustomize.




  • Reliability engineering and incident response. Help build and run incident management processes: severity definitions, escalation paths, on-call practices. Lead incident response, debugging, and root-cause analysis. Write postmortems and drive systemic reliability improvements from what they surface. Improve observability across metrics, logging, tracing, and dashboards. Support GPU and batch workloads running on Kubernetes.




  • Security and access management. Provide security-conscious feedback on platform architecture decisions. Own cloud IAM governance: roles, policies, and access boundaries across accounts and services. Lead compliance-adjacent work including audit-readiness, partner certification requirements, and supporting responses to customer security questionnaires.



  • Platform tooling and developer experience. Improve CI/CD pipelines and infrastructure validation. Support engineers with infrastructure debugging, environment setup, and performance issues. Contribute to tooling and automation in Python and Bash. Take on adjacent responsibilities as needed in a startup environment.


Required Qualifications



  • Experience. 5+ years in SRE, DevOps, or infrastructure engineering roles, with a track record of operating production systems across multiple regions.




  • Terraform. Modules, state management, and multi-environment patterns.




  • AWS depth. Solid experience across VPC, IAM, EKS, S3, and CloudWatch.




  • Kubernetes expertise. Cluster operations, autoscaling, RBAC, and Helm.




  • CI/CD and GitOps. Experience with GitHub Actions, ArgoCD, or similar workflows.




  • Networking fundamentals. CIDR, DNS, load balancing, VPN, and cross-region connectivity.




  • Observability. Experience with tooling such as Prometheus and Grafana.




  • Scripting. Comfort with Python and Bash for tooling and automation.




  • Cross-platform familiarity. Working knowledge of both Linux and Windows environments. Operational experience supporting Windows-based workloads is a meaningful advantage.



  • Pragmatism and ownership. Comfortable in a fast-moving startup with evolving priorities. You take ownership of systems while collaborating closely with other teams, and you're pragmatic about tradeoffs between speed, reliability, and complexity.


Preferred Qualifications



  • Windows on Kubernetes. Experience with Windows node pools, Windows AMIs, and GPU-adjacent components on K8s.




  • GPU scheduling. Familiarity with GPU scheduling on Kubernetes, including NVIDIA device plugin configuration.




  • Domain workloads. Experience supporting simulation, ML, or rendering workloads in cloud infrastructure.




  • AWS extras. Exposure to AWS Storage Gateway, Active Directory integrations, or AWS Transfer Family.




  • Service mesh. Familiarity with service proxy or service mesh patterns.




  • Container OS. Experience with container-optimized OS images (e.g., Bottlerocket, Packer).



  • Cost optimization. Cloud cost optimization at scale.


Core Tools

Terraform · AWS · Kubernetes · Helm · Kustomize · ArgoCD · GitHub Actions · Prometheus · Grafana · Docker · Python · Bash

What Makes a Great Candidate

You think in failure modes and proactively surface issues. You hold a principled view on security and push back constructively when designs introduce unnecessary risk. You communicate clearly across engineering, product, and customer-facing teams, flagging issues with urgency proportional to customer impact. You take end-to-end ownership of complex efforts and know when to push for the clean solution versus the pragmatic one.Base salary range of CAD $145,000–$185,000, depending on skills, qualifications, and experience, plus equity, full health/dental/vision coverage, learning stipend, and generous vacation. This role is remote-friendly across Canada and the US Pacific Northwest.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Créer une alerte emploi pour cette recherche

Remote Senior Site Reliability Engineer • Markham, Ontario

Offres similaires

Senior Site Reliability Engineer

Guidewire Softwaretoronto, on, Canada
Temps plein

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

CapgeminiToronto, ON, CA
Temps plein

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

MorningstarToronto, ON, CA
Temps plein

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Voir plus

 • Offre sponsorisée

Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)

MongoDBtoronto, on, Canada
Temps plein

MongoDB’s Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture.This relatively new team is bu... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

DexianToronto, ON, CA
Temps plein

Working Location: Toronto, ON [Hybrid 2 days a week in office].The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Voir plus

 • Offre sponsorisée

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Télétravail
Temps plein

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Temps plein

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Temps plein

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

Future Secure AIToronto, ON, CA
Temps plein

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer (Remote-First)

VySystemsToronto, ON, CA
Télétravail
Temps plein

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, ON, CA
Temps plein

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

Magnet Forensicstoronto, on, Canada
Temps plein

Magnet Forensics is a global leader in the development of digital investigative software that acquires, analyzes, and shares evidence from computers, smartphones, tablets, and IoT-related devices.O... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

RootlyToronto, ON, CA
Temps plein

At Rootly, we are on a mission to be the go‑to way companies respond when things go wrong, helping every organization be more reliable.We do this by building an industry‑leading incident management... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

Momentum Financial Services GroupToronto, ON, CA
Temps plein

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsToronto, ON, CA
Temps plein

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

Socket.devToronto, ON, CA
Temps plein

We are seeking a Senior Consultant in Site Reliability Engineering (Network SRE) to lead network-centric reliability practices across the Shared Platform ecosystem.This role focuses on ensuring res... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Télétravail
Temps plein

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

iManageToronto, ON, CA
Temps plein

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams – SRE teams are anchored t... Voir plus

 • Offre sponsorisée

Site Reliability Engineer - Canada Wide - Remote

NewtonToronto, ON, CA
Télétravail
Temps plein

Say hello to Newton! We're changing how Canadians trade crypto.Our goal? To make financial freedom something everyone can achieve.We give our customers the tools and knowledge they need to navigate... Voir plus

 • Offre sponsorisée

Site Reliability Engineer (Senior or Staff), Deployments

AlleyCorptoronto, on, Canada
Temps plein

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Voir plus