Talent.com

Reliability engineer Jobs in Surrey, BC

Create a job alert for this search

Reliability engineer • surrey bc

Last updated: 5 hours ago
  • New!
Founding SRE Engineer (Infrastructure & Reliability)

Founding SRE Engineer (Infrastructure & Reliability)

OpusBurnaby, Metro Vancouver Regional District, CA
Full-time
AI video agent, built for authenticity on social media.We envision a world where everyone can authentically share their story through video, with no expertise needed. Within just 18 months of our la...Show moreLast updated: 5 hours ago
Maintenance Supervisor - Plant Reliability & TPM Leader

Maintenance Supervisor - Plant Reliability & TPM Leader

PepsiCo Deutschland GmbHDelta, Metro Vancouver Regional District, CA
Full-time
A global beverage and food company is seeking a Maintenance Supervisor in Delta, Canada.This role involves overseeing preventative maintenance and repairs of plant equipment.Ideal candidates will h...Show moreLast updated: 2 days ago
Senior DevOps Engineer - ML Infrastructure (Remote)

Senior DevOps Engineer - ML Infrastructure (Remote)

Serve RoboticsSurrey, Metro Vancouver Regional District, CA
Remote
Full-time
A leading technology company is seeking a Senior DevOps Engineer to deploy and maintain machine learning infrastructure.Candidates should have over 5 years of experience in DevOps or Infrastructure...Show moreLast updated: 30+ days ago
Mechanical Engineer III

Mechanical Engineer III

VerathonBurnaby, Metro Vancouver Regional District, CA
Full-time
The company has significantly impacted patient care in bladder volume measurement and airway management, becoming the market leader in each area. Its BladderScan portable ultrasound and GlideScope v...Show moreLast updated: 30+ days ago
Maintenance Manager : Lead Reliability & Uptime

Maintenance Manager : Lead Reliability & Uptime

Cobalt SearchCity of Langley, Metro Vancouver Regional District, CA
Full-time
A prominent manufacturing company located in Langley, British Columbia is looking for a Maintenance Manager to lead a multifaceted maintenance team. This role emphasizes leadership, problem-solving,...Show moreLast updated: 26 days ago
Senior Software Development Engineer, ML Platform

Senior Software Development Engineer, ML Platform

Remitly Inc.Burnaby, Metro Vancouver Regional District, CA
Full-time
Senior Software Development Engineer, ML Platform page is loaded## Senior Software Development Engineer, ML Platformlocations : Burnaby, British Columbia, Canadatime type : Full timeposted on : ...Show moreLast updated: 30+ days ago
Senior Software Engineer - Changing the face of sports

Senior Software Engineer - Changing the face of sports

Uplifter Inc.Burnaby, Metro Vancouver Regional District, CA
Full-time
Senior Software Engineer - Changing the face of sports at.Senior Software Engineer - Changing the face of sports.Hybrid – Burnaby (Vancouver), BC. Develop and maintain backend services using Python / ...Show moreLast updated: 30+ days ago
  • Promoted
Senior Site Reliability Engineer

Senior Site Reliability Engineer

Targeted TalentBurnaby, BC, Canada
Permanent
We are looking for an experienced.Senior Site Reliability Engineer.Our client is a global enterprise company with a product that you've likely used. Experience with coding / software development, ...Show moreLast updated: 30+ days ago
Mechanical Engineer

Mechanical Engineer

TPD® Workforce SolutionsSurrey, Metro Vancouver Regional District, CA
Full-time
Get AI-powered advice on this job and more exclusive features.Direct message the job poster from TPD® Workforce Solutions. Workforce Manager @ TPD® | Recruiter for Engineers & Geoscientists BC | Dir...Show moreLast updated: 16 days ago
Refinery Fixed Equipment Inspector : Reliability Focus

Refinery Fixed Equipment Inspector : Reliability Focus

Parkland CorporationBurnaby, Metro Vancouver Regional District, CA
Full-time
A leading energy company in Burnaby is seeking a Fixed Equipment Integrity Inspector to ensure the reliability and integrity of the refinery's fixed equipment assets. The successful candidate will c...Show moreLast updated: 8 days ago
Maintenance Manager, 24 / 7 Plant Reliability Leader

Maintenance Manager, 24 / 7 Plant Reliability Leader

Veolia Water Technologies & SolutionsBurnaby, Metro Vancouver Regional District, CA
Full-time
A leading environmental services company in Burnaby is looking for a Maintenance Manager to direct maintenance strategies and lead an in-house team. The role involves managing schedules, budgets, an...Show moreLast updated: 2 days ago
Senior Backend Platform Engineer — Data & Trust

Senior Backend Platform Engineer — Data & Trust

RemitlyNew Westminster, BC, CA
Full-time
A financial technology company in New Westminster is seeking an experienced Software Engineer for their Customer Data Platform team. You will build and manage services that handle sensitive customer...Show moreLast updated: 30+ days ago
Snowflake Platform Engineer

Snowflake Platform Engineer

Best BuyBurnaby, Metro Vancouver Regional District, CA
Full-time
Snowflake Platform Engineer page is loaded## Snowflake Platform Engineerremote type : Remotelocations : 00000 Canadian Headquarterstime type : Full timeposted on : Posted Yesterdayjob requisiti...Show moreLast updated: 8 days ago
Hybrid Infrastructure Engineer — Cloud, IaC & Security

Hybrid Infrastructure Engineer — Cloud, IaC & Security

PeopleToGo Inc.Burnaby, Metro Vancouver Regional District, CA
Full-time
A leading technology firm is seeking an experienced Infrastructure Engineer in Burnaby.You will design, build, and operate secure and scalable infrastructure platforms while managing critical cloud...Show moreLast updated: 2 days ago
Warehouse Maintenance & Reliability Manager

Warehouse Maintenance & Reliability Manager

DB SchenkerDelta, Metro Vancouver Regional District, CA
Full-time
A global logistics company in Delta, Canada is seeking a Facilities Maintenance Supervisor to ensure the operational efficiency of their refrigerated warehouse. The role involves directing maintenan...Show moreLast updated: 1 day ago
Maintenance Supervisor - Equipment Reliability & Leadership

Maintenance Supervisor - Equipment Reliability & Leadership

Kruger Products | Produits KrugerNew Westminster, BC, CA
Full-time
A leading manufacturer in tissue products is seeking a Maintenance Supervisor in New Westminster, BC.This role involves supervising a skilled team in maintenance activities, ensuring production rel...Show moreLast updated: 1 day ago
Maintenance & Reliability Co-op — Engineering Student

Maintenance & Reliability Co-op — Engineering Student

Trans MountainBurnaby, Metro Vancouver Regional District, CA
Full-time
A leading pipeline company in Burnaby is seeking a Maintenance and Reliability Co‑op Student for an 8-month term starting in May 2026. This role involves supporting maintenance planning, coordinatin...Show moreLast updated: 26 days ago
Maintenance & Reliability Co-op — Engineering Student

Maintenance & Reliability Co-op — Engineering Student

Trans Mountain CorporationBurnaby, Metro Vancouver Regional District, CA
Full-time
A leading pipeline company is seeking a Maintenance and Reliability Co-op Student for an 8-month term based in Burnaby, British Columbia. The role involves supporting maintenance planning, organizin...Show moreLast updated: 26 days ago
Software Engineer - Contingent

Software Engineer - Contingent

Ritchie Bros.Burnaby, Metro Vancouver Regional District, CA
Full-time
Software Engineer - Contingent role at Ritchie Bros.The Software Engineer is responsible for analysis, development and ongoing support of key internal software systems at Ritchie Bros.Salesforce an...Show moreLast updated: 30+ days ago
People also ask
Founding SRE Engineer (Infrastructure & Reliability)

Founding SRE Engineer (Infrastructure & Reliability)

OpusBurnaby, Metro Vancouver Regional District, CA
5 hours ago
Job type
  • Full-time
Job description

🎨 OpusClip is the world's No.1 AI video agent, built for authenticity on social media.

We envision a world where everyone can authentically share their story through video, with no expertise needed. Within just 18 months of our launch, over 10 million creators and businesses have used OpusClip to enhance their social presence.

We have raised $50 million in total funding and are fortunate to have some of the most supportive investors, including SoftBank Vision Fund, DCM Ventures, Millennium New Horizons, Fellows Fund, AI Grant, Jason Lemkin (SaaStr), Samsung Next, GTMfund, Alumni Ventures, and many more.

Check out our latest coverage by Business Insider featuring our product and funding milestones, and our recognition as one of The Information's 50 Most Promising Startups in 2024.

Headquartered in Palo Alto, we are a team of 100 passionate and experienced AI enthusiasts and video experts, driven by our core values :

Be a Champion Team

Prioritize Ruthlessly

Ship fast, Quality Follows

Obsess over customers

Be a part of this exciting journey with us!

The Mission

We are looking for a hands-on Founding SRE to own the reliability and scalability of the OpusClip platform. You will stabilize our processing clusters, design isolated environments for our largest Enterprise partners, and serve as the technical bridge between infrastructure and our 15M+ users.

You will engineer the infrastructure strategy that underpins our trust and reliability in the market. You will help set up oncall rotation and own the full incident lifecycle from minimizing Time-to-Detect to tracking post mortem actions ensuring that our high-velocity growth never compromises our performance.

Key Responsibilities

Infrastructure Architecture & Cluster Operations

Architect Dedicated Environments : Lead the design and implementation of high-throughput, isolated processing clusters for Enterprise clients. You will build the "paved road" to ensure strict High Availability (HA) without noisy neighbor interference.

Scale Production : Drive general improvements in our Temporal clusters and production Kubernetes environments. You will operationalize scaling strategies that support both self-serve consumers and high-touch Enterprise contracts.

Technical Execution : Be hands-on with the stack to optimize resource allocation, reduce latency, and enforce isolation strategies for critical accounts.

Monitoring, Alerting & Detectability

Beat the Customer to the Alert : Overhaul our Datadog observability suite to aggressively reduce Time-to-Detect (TTD) . You ensure we identify latency spikes and stalled projects before users do.

Threshold Tuning : tune alert thresholds to eliminate noise and focus on "symptom-based" alerts that reflect the actual user experience.

External SLO Ownership : Define and report on Service Level Objectives (SLOs), acting as the internal guarantor that we are meeting the targets we sold.

Incident Command & "Extreme Ownership"

First Responder & Mitigation : Serve as the first line of defense during outages. You will own immediate mitigation, including cluster debugging and manual scaling intervention if Horizontal Pod Autoscalers (HPA) fail or lag.

Drive Recovery Metrics : You are accountable for shortening Time-to-Mitigation (TTM) and Time-to-Recover (TTR) . Your priority is to stop the bleeding first, then fix the wound.

Root Cause Analysis : Lead the post-mortem process to determine Time-to-Root Cause and implement systemic fixes. You will translate these technical findings into clear updates for Customer Experience (CX) and Leadership.

Accountability : Work collaboratively with Engineering Owners to track improvements against the reliability roadmap. You are responsible for flagging risks early and resetting expectations on platform performance when necessary.

Cross-Functional Bridge : Serve as the primary technical voice to the Customer Experience (CX), Sales, and Leadership teams. You will translate technical constraints and roadmaps into clear, honest updates for stakeholders.

Qualifications

Production K8s & Temporal : Expert-level ability to debug Kubernetes internals (HPA logic, node scaling) and operate stateful workflow engines (Temporal) at scale.

Incident Command : Proven track record as a primary first responder, demonstrating the ability to aggressively reduce Time-to-Mitigation (TTM) and Time-to-Recover (TTR) .

Observability Architecture : Experience architecting Datadog SLOs and tuning alerts to distinguish system noise from actual user pain.

Automation : Strong proficiency in Python or Bash to automate manual recovery and operational tasks.

⭐️ Bonus : Experience scaling GPU / video rendering workloads or thriving in early-stage, high-velocity startups.

The Tech Stack

Orchestration & Compute : Kubernetes (GKE), Docker, Horizontal Pod Autoscaling (HPA).

Workflow Engine : Temporal

Observability : Datadog (APM, Custom Metrics, Alerting).

Infrastructure as Code : Terraform or similar IaC tools.

Scripting & Backend : Python (primary), Bash.

Data & Storage : Redis, Milvus (Vector DB), Postgres.

EEO

OpusClip is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristics. OpusClip considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Opus Clip is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures.

#J-18808-Ljbffr