Talent.com
Cerebras
Staff Site Reliability Engineer – Automation and PlatformCerebras • Toronto, Ontario, CA
Staff Site Reliability Engineer – Automation and Platform

Staff Site Reliability Engineer – Automation and Platform

Cerebras • Toronto, Ontario, CA
30+ days ago
Job type
  • Full-time
Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real‑time iteration and increasing intelligence via additional agentic computation.

Cerebras works with the leading model labs, global enterprises, and cutting‑edge AI‑native startups. OpenAI recently announced a multi‑year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high‑speed inference.

About The Role

We are building a high‑performance SRE function to support one of the world’s fastest‑growing AI inference services, powered by the Wafer‑Scale Engine (WSE). This team will help deliver world‑class, ultra‑reliable inference infrastructure for leading model builders such as OpenAI and other frontier labs.

As a Staff SRE, you will lead the engineering effort to eliminate toil at scale by driving implementation of self‑service delivery pipelines, shared observability common tooling. This role starts with ~1 month of hands‑on operational immersion to gain deep familiarity with our current stack, production pain points, and high‑stakes workflows.

From there, your primary focus shifts to architecting and delivering the "tomorrow" layer: declarative GitOps‑driven CD for model releases, capacity provisioning and cluster upgrades. Success over the first year in this role will be defined by enabling core teams, product managers, external customers, and cluster stakeholders to operate in a fully self‑service model with strong reliability guarantees.

You will partner with our early‑career SRE sub‑team, who own day‑to‑day operations. This will allow you to deeply understand their pain points, automate their toil, and mentor them as platform engineers.

You will collaborate with the tech leads and the leadership team across core, cluster, cloud, and product stakeholders. This work will shift reliability from an ops‑only burden to a shared engineering discipline that underpins frontier AI inference at scale.

If you are a proven Staff+ engineer who enjoys turning complexity into elegant reliability at scale, this is your chance to lead this transformation from the front.

This role does not require 24/7 on‑call rotations.

Key Responsibilities

  • Define and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud‑based solutions.
  • Architect self‑service platforms and internal tooling that lets product teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.
  • Define and evolve reliability practices for inference workloads, including SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi‑datacenter and on‑prem environments.
  • Mentor mid‑level SREs, support critical incident escalations, and use production pain points to prioritize the highest‑leverage automation work.
  • Measure and drive impact through clear metrics, including toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self‑service workflows.

Required Experience & Skills

  • 8+ years in SRE, infrastructure engineering, or platform engineering, with a strong record of improving automation and reliability at large scale in FAANG, hyperscaler, or similarly demanding environments.
  • Deep expertise operating large scale heterogenous clusters with a proprietary cloud control plane.
  • Proven track record designing and delivering CI/CD or GitOps systems using Argo CD or similar tools, with strong safety and observability built in.
  • Hands‑on experience with observability systems such as Loki, Tempo, Mimir, and Prometheus.
  • Ability to lead complex projects end to end, influence cross‑functional stakeholders, and communicate technical direction clearly.

Nice‑to‑Haves

  • Experience with Bazel or other large‑scale build systems in production.
  • Background in AI/ML inference systems, including model serving runtimes, GPU or wafer‑scale orchestration, latency and accuracy SLOs, or drift monitoring.
  • Prior work on predictive autoscaling, chaos engineering, or cost‑aware capacity planning for compute‑intensive workloads.

Location

  • SF Bay Area
  • Toronto

Why Join Cerebras

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting‑edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non‑corporate work culture that respects individual beliefs.

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third‑party tools process personal data. For more details, click here to review our CCPA disclosure notice.

#J-18808-Ljbffr
Create a job alert for this search

Staff Site Reliability Engineer – Automation and Platform • Toronto, Ontario, CA

Similar jobs

Senior Site Reliability Engineer - Automation & Uptime

CapgeminiToronto, ON, CA
Full-time

A global technology consulting firm is seeking a Talent Acquisition Business Partner to drive hiring initiatives in Toronto, Canada.This mid-senior level full-time role involves implementing monito... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, Ontario, Canada
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Staff Site Reliability Engineer - $153,400 - $220,400 A Year - Remote

CoalitionEast York, Canada
Remote
Full-time

Seeking a Staff Site Reliability Engineer to lead AI enablement initiatives for software development, focusing on reliability, security, and tooling.This role involves designing, building, and adop... Show more

 • Promoted

Site Reliability Engineer

Socket.devtoronto, on, Canada
Full-time

We are seeking a Senior Consultant in Site Reliability Engineering (Network SRE) to lead network-centric reliability practices across the Shared Platform ecosystem.This role focuses on ensuring res... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, ON, CA
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office].The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Show more

 • Promoted

Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc.Toronto, ON, CA
Full-time

A leading supply chain solutions provider is seeking a Site Reliability Engineer to optimize and ensure the reliability of their cloud infrastructure across AWS and Kubernetes.This role emphasizes ... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Site Reliability Engineer

Future Secure AIToronto, ON, CA
Full-time

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Show more

 • Promoted

Staff Site Reliability Engineer – Automation and Platform

CerebrasToronto, ON, CA
Full-time

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than... Show more

 • Promoted

Site Reliability Engineer (Senior Or Staff), Deployments

AlleyCorpToronto, Canada
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Lead Site Reliability Engineer - C$140,000 - C$178,000 A Year

EpamEast York, Canada
Full-time

Lead Site Reliability Engineer to ensure system reliability, observability, and performance monitoring for digital trading products.Responsible for leading monitoring initiatives, defining reliabil... Show more

 • Promoted

Site Reliability Engineer, Aiops - C$94,000 - C$110,000 A Year

CognizantToronto County, Canada
Full-time

Site Reliability Engineer implementing AI-driven observability pipelines and automation for production environments to improve system reliability. Show more

 • Promoted

PheedLoop Senior Site Reliability Engineer

PheedLoopToronto, ON, CA
Full-time

PheedLoop is seeking a Senior Site Reliability Engineer to maintain and enhance the tech behind live events.This pivotal role prioritizes system reliability and performance.You'll be tasked with de... Show more

 • Promoted

Staff Site Reliability Engineer - Confluent Incident Management & Reliability

IBMToronto, ON, CA
Full-time

Your Role and Responsibilities.Confluent Cloud processes millions of events per second across AWS, GCP, and Azure.When incidents happen in a multi‑cloud streaming platform, they happen at scale—dat... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, Ontario, Canada
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer - C$110,000 - C$120,000 A Year

Momentum Financial Services GroupNorth York, Canada
Full-time

Site Reliability Engineer needed to ensure availability, performance, and resilience of financial platforms by automating operations, defining SLOs, and engineering systems for failure recovery.Wor... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Sr. Site Reliability Engineer I - C$184,088 - C$294,540 A Year - Remote

AxonToronto, Canada
Remote
Full-time

Join Axon and be a Force for Good.At Axon, we're on a mission to Protect Life.We're explorers, pursuing society's most critical safety and justice issues with our ecosystem of devices a... Show more

 • Promoted

Site Reliability Engineer (Senior or Staff), Deployments

AlleyCorptoronto, on, Canada
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Staff Site Reliability Engineer - C$124,200 - C$166,700 A Year

Walt Disney Animation StudiosEast York, Canada
Full-time

Seeking a Staff Site Reliability Engineer with expertise in Linux, software development (Python, Go, Java, Node), CI/CD tools, Git, cloud platforms (AWS, GCP, Azure), and container technologies (Do... Show more