Talent.com
Baseten
Site Reliability EngineerBaseten • Montreal (administrative region), Quebec, CA
Site Reliability Engineer

Site Reliability Engineer

Baseten • Montreal (administrative region), Quebec, CA
20 days ago
Salary
CA$165,000.00 yearly
Job type
  • Full-time
Job description

About Baseten

Baseten powers mission‑critical inference for the world’s most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting‑edge models into production. We’re growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

As a Site Reliability Engineer at Baseten, you’ll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You’ll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently. You’ll work closely with engineering, forward‑deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company.

EXAMPLE INITIATIVES

You’ll work on projects like these as part of the SRE team:

  • Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services.
  • Building AI‑assisted tooling for incident triage and response.

Responsibilities

  • Own the reliability of Baseten’s multi‑cloud Kubernetes infrastructure, including incident response, post‑mortems, and remediation tracking.
  • Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code.
  • Author, validate, and improve runbooks for recurring failure patterns, ensuring they’re structured for low‑context, safe execution.
  • Identify high‑frequency failure patterns and convert them into automated mitigations or self‑healing automations.
  • Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management.
  • Define and instrument SLOs and SLIs across customer workloads and internal services.
  • Navigate ambiguity, make principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define.

Requirements

  • Extensive hands‑on experience with Kubernetes (multi‑cloud experience across EKS, GKE, or similar is a strong plus).
  • Experience in building and maintaining scalable infrastructure.
  • Strong foundation in observability tooling: metrics (VictoriaMetrics, Prometheus), logging (Loki, ELK), dashboards (Grafana), and alerting pipelines. Observability‑as‑code experience is a plus.
  • Experience with infrastructure‑as‑code (Terraform, Helm) and GitOps workflows (Flux CD, ArgoCD).
  • Experience writing and improving runbooks, leading incident response, and doing post‑mortem analysis.
  • Comfort working at the intersection of engineering and operations — you write code, but you also think deeply about process, escalation paths, and operational leverage.
  • Familiarity with incident management platforms (incident.io or similar) is a plus.
  • No prior ML experience required, but curiosity about how ML models are deployed and served at scale will serve you well.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents.
  • Flexible PTO policy including company‑wide Winter Break (our offices are closed from Christmas Eve to New Year’s Day!).
  • Paid parental leave.
  • Fertility and family‑building stipend through Carrot.
  • Company‑facilitated 401(k).
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now

to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward‑thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Compensation Range: $165K - $330K

#J-18808-Ljbffr
Create a job alert for this search

Site Reliability Engineer • Montreal (administrative region), Quebec, CA

Similar jobs

Site Reliability Engineer

Vertex Elite LLCRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Duration: ContractKey Skills:Monitoring / Observability tools - Dynatrace, ELK etc.Platform/ cloud Observability - OpenShift, Prometheus / Azure Cloud etc.Key Responsibilities:Collaborate with vari... Show more

 • Promoted

Radformation Senior Engineer - Adaptive Systems

Radformationmontreal (administrative region), qc, Canada
Full-time

Elevate your career as a Senior Engineer in Adaptive Systems at Radformation, working remotely to transform cancer care.Build software that enhances treatment efficiency for radiation oncology.The ... Show more

 • Promoted

Sr. Engineer

TechDoQuestmontreal (administrative region), qc, Canada
Full-time

Perform icing numerical simulations on complex aerodynamic configurations.Prepare, execute, and analyze high‑lift and icing wind tunnel test campaigns, including CFD and certification.Architect and... Show more

 • Promoted

Senior Full-Stack Engineer - Accessibility & Inclusive Tech (Remote)

Accessibility Partners CanadaMontreal (administrative region), QC, CA
Remote
Full-time

A leader in accessible technology is seeking a Senior Full-Stack Developer to create equitable and accessible digital systems.This role involves developing both front-end and back-end systems, focu... Show more

 • Promoted

Site Reliability Engineer

Hunter BondMontréal, Canada
Full-time

Role: DevOps EngineerClient: Most Elite Tech Firm in CanadaCompensation: Up to $200k CAD + Bonus + PackageLocation: MontrealOverviewAn Elite FinTech Firm is looking for a highly talented DevOps Eng... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsMontreal (administrative region), QC, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Specialist Site Reliability Engineer

Global Talent Alliance, CanadaMontreal
Full-time

About the job Specialist Site Reliability Engineer.The role of the Specialist Site Reliability Engineer (SRE) is to execute RAM analysis and engineering in support of the I&T solutions.The overall ... Show more

 • Promoted

Senior Ii Site Reliability Engineer

Akamai TechnologiesRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Job DescriptionJoin our SRE team! Our team uses large datasets to analyze and measure the performance and reliability of our platform.We are networking data scientists: we combine our knowledge of ... Show more

 • Promoted

Senior Tailings Engineer - Remote Leadership & Design

StantecMontreal (administrative region), QC, CA
Remote
Full-time

A leading engineering firm is seeking an experienced engineer specializing in mine tailings management to join their Montreal team.This pivotal role involves leading complex technical projects, dev... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificMontreal (administrative region), QC, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Site Reliability Engineer

BasetenMontréal, Canada
Full-time

About BasetenBaseten powers mission‐critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting applied AI resear... Show more

 • Promoted

SAR Payload Systems Engineering Advisor

Sky Systems, Inc. (SkySys)montreal (administrative region), qc, Canada
Full-time

SAR Payload Systems Engineering Advisor.Duration: 12 months – 40 hours per week.Location: Montreal – West Island, 4 days per week in office.Salary: CAD $125K - $140K Annually with Standard Benefits... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Site Reliability Engineer

SecurityScorecardmontreal (administrative region), qc, Canada
Full-time

SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries.Founded in 2013 by security and risk experts Dr.Alex Ya... Show more

 • Promoted

Site Supervisor

Nasittuq CorporationMontreal (administrative region), QC, CA
Full-time

Join Nasittuq for a unique and rewarding experience!.Nasittuq Corporation (from the Inuktitut word meaning “looking out from the highest point”) operates and maintains the North Warning System (NWS... Show more

 • Promoted

Site Reliability Engineer

MaintainXMontreal
Full-time

MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 13,000+ customers including Duracell, Shell, Cintas, and Brenntag.We raised $150M in Series D funding ... Show more

 • Promoted

Site Reliability Engineer - Tech Talent International

Tech Talent InternationalMontreal
Full-time

Join Tech Talent International as a Senior Site Reliability Engineer, specializing in Automation & Observability, located in Montreal.This hybrid role focuses on enhancing production efficiency and... Show more

 • Promoted

Site Reliability Engineer (Linux / Cloud Infrastructure)

Atlantis IT GroupMontréal, Quebec, Canada
Full-time

Site Reliability Engineer (Linux / Cloud Infrastructure) role with hands-on experience across Linux, distributed systems, scripting, databases, monitoring, containers, cloud SaaS integrations, mess... Show more