Deployment Engineer, AI InferenceCerebras Systems Inc. • Toronto, Canada

Deployment Engineer, AI Inference

Cerebras Systems Inc. • Toronto, Canada

24 days ago

Job type

Full-time

Job description

About Cerebras Systems

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer‑scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of one device. This enables industry‑leading training and inference speeds and lets machine learning users run large‑scale ML applications without the hassle of managing hundreds of GPUs or TPUs.

Our customers include global corporations, national labs and top‑tier healthcare systems. In 2024 we launched Cerebras Inference, the fastest generative AI inference solution, over 10 times faster than GPU‑based hyperscale cloud inference services.

About The Role

We are seeking a highly skilled Deployment Engineer to build and operate cutting‑edge inference clusters on the world’s largest computer chip, the Wafer‑Scale Engine (WSE). You will play a critical role in ensuring reliable, efficient, and scalable deployment of AI inference workloads across our global infrastructure. On the operational side, you’ll own the rollout of new software versions, AI replica updates, and capacity reallocations across our custom‑built, high‑capacity datacenters. Beyond operations, you’ll drive improvements to telemetry, observability, and fully automated pipelines using advanced allocation strategies to maximize utilization of large‑scale computer fleets.

The ideal candidate combines hands‑on operation rigor with strong systems engineering skills and thrives on building resilient pipelines that keep pace with cutting‑edge AI models.

This role does not require 24 / 7 hour on‑call rotations.

Responsibilities

Deploy AI inference replicas and cluster software across multiple datacenters

Maximize capacity allocation and optimize replica placement using constraint‑solver algorithms

Operate bare‑metal inference infrastructure while supporting transition to K8S‑based platform

Develop and extend telemetry, observability and alerting solutions to ensure deployment reliability at scale

Develop and extend a fully automated deployment pipeline to support fast software updates and capacity reallocation at scale

Translate technical and customer needs into actionable requirements for the Dev Infra, Cluster, Platform and Core teams

Stay up to date with the latest advancements in AI compute infrastructure and related technologies

Skills And Requirements

2–5 years of experience operating on‑prem compute infrastructure (ideally in Machine Learning or High‑Performance Compute) or developing and managing complex AWS‑based infrastructure for hybrid deployments

Strong proficiency in Python for automation, orchestration, and deployment tooling

Solid understanding of Linux‑based systems and command‑line tools

Extensive knowledge of Docker containers and container orchestration platforms like K8S

Familiarity with spine‑leaf (Clos) networking architecture

Proficiency with telemetry and observability stacks such as Prometheus, InfluxDB and Grafana

Strong ownership mindset and accountability for complex deployments

Ability to work effectively in a fast‑paced environment

Location

SF Bay Area

Toronto

Why Join Cerebras

Build a breakthrough AI platform beyond the constraints of the GPU

Publish and open‑source cutting‑edge AI research

Work on one of the fastest AI supercomputers in the world

Enjoy job stability with startup vitality

Our simple, non‑corporate work culture respects individual beliefs

Apply Today and Become Part of the Forefront of Groundbreaking Advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. Inclusive teams build better products and companies, and we empower people to do their best work through continuous learning, growth, and support of those around them.

#J-18808-Ljbffr

Create a job alert for this search

Deployment Engineer AI Inference • Toronto, Canada

Similar jobs

AI Engineer

TheAppLabb • Toronto, ON, Canada

Full-time

AI-powered digital solutions, mobile app development, and emerging technologies.We leverage data-driven insights to enhance digital experiences and drive business growth. We are looking for a Machin...Show more

Last updated: 30+ days ago • Promoted

Remote AI / ML Solutions Architect — AWS & GenAI Expert

Avahi • Toronto C6A, ON, Canada

Remote

Full-time

A cloud-first consulting company is looking for an experienced AI / ML Solutions Architect to join their remote-first team. The successful candidate will leverage AWS services to design scalable AI / ML...Show more

Last updated: 30+ days ago • Promoted

Senior LLMOps Engineer — Cloud AI Inference + Equity

TEEMA Solutions Group • Toronto, Canada

Full-time

A rapid-growth technology firm in Toronto is seeking a Staff LLMOps Engineer to lead the design and optimization of large language model infrastructure on the cloud. The ideal candidate has over 6 y...Show more

Last updated: 16 days ago • Promoted

Strategic Deployment Lead for Enterprise AI (Toronto)

Peregrine • Toronto

Full-time

A technology solutions provider in Toronto seeks a Deployment Strategist to lead the product deployment and enhance customer impact. Ideal candidates will have 2–5 years of experience in deploying e...Show more

Last updated: 17 days ago • Promoted

Lead AI Engineer

Harnham • Toronto, ON, Canada

Full-time

Toronto, ON - 3 days onsite / week.CAD + bonus + LTI; 300,000 - 400,000 CAD total.Harnham is partnering with one of the most well known financial services companies, which is looking for an experienc...Show more

Last updated: 6 days ago • Promoted

Forward Deployed Engineer, Agentic AI Platform

Cohere • Toronto, Canada

Full-time

A technology company based in Toronto is looking for an Entry-Level Forward Deployed Engineer to help develop its AI workspace platform, North. This unique role involves building features and deploy...Show more

Last updated: 24 days ago • Promoted

AI Inference Digital Design Engineer

Taalas • Toronto, Canada

Full-time

A tech company is seeking a passionate Digital Design Engineer in Toronto to work on complex AI Inference challenges.Candidates should have a degree in Electrical or Computer Engineering and profic...Show more

Last updated: 30+ days ago • Promoted

AI Engineer

GEI Consultants • Markham, ON, Canada

Full-time

The AI Engineer is responsible for the development of AI solutions, typically leveraging pretrained models and copilots, to support GEI’s priority digital and AI initiatives.This role focuses...Show more

Last updated: 30+ days ago • Promoted

Cloud AI Inference Engineer - Golang

Huawei Canada • Markham

Full-time +1

A technology firm in York Region, Canada is seeking an Engineer for a 12-month contract.The role includes collaborating with senior engineers to design AI inference frameworks using Golang, optimiz...Show more

Last updated: 11 hours ago • Promoted • New!

AI Platform Engineer : Build & Evaluate AI Systems

Armilla AI • Toronto, Canada

Full-time

A leading AI-focused startup in Toronto is seeking an experienced AI Engineer to shape the future of AI risk management.In this pivotal role, you will be responsible for building core AI tools and ...Show more

Last updated: 30+ days ago • Promoted

Senior AI Platform Engineer — Remote

Dayforce US, Inc. • Toronto, Canada

Remote

Full-time

A global HCM company is seeking an AI Platform Developer Senior to build scalable AI application platforms.This role focuses on designing infrastructure and developing tools to enhance developer wo...Show more

Last updated: 22 days ago • Promoted

Sr. Deployment Engineer, AI Inference

Cerebras • Toronto

Full-time

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programm...Show more

Last updated: 11 hours ago • Promoted • New!

Forward-Deployed AI Engineer for Enterprise Solutions

Google Inc. • Toronto, Canada

Full-time

A leading technology company is seeking a Forward Deployed Engineer to utilize Generative AI in solving customer challenges. The role involves software development, deployment, and collaboration wit...Show more

Last updated: 19 days ago • Promoted

AI Systems & Inference Frameworks Engineer

adaption • Toronto, Canada

Full-time

About Us Most AI is frozen in place - it doesn't adapt to the world.Our mandate is to build efficient intelligence that evolves in real‑time. Our vision is AI systems that are flexible, personalized...Show more

Last updated: 22 days ago • Promoted

Remote-First LLM AI Research Engineer (NLP Innovator)

Upfeat Media, Inc. • Toronto, Canada

Remote

Full-time

A leading company in applied AI seeks a Large Language Model AI Research Engineer passionate about NLP and innovative solutions. This role involves developing datasets, engaging with cutting-edge te...Show more

Last updated: 24 days ago • Promoted

AI Solutions Engineer - Lead Deployment & Impact

Inbenta • Toronto, Canada

Full-time

A global technology company in Toronto is seeking an AI Solutions Engineer to lead technical engagement across the customer lifecycle. The role involves solution design, implementation, and ongoing ...Show more

Last updated: 22 days ago • Promoted

Senior AI Inference Systems Engineer — Equity Eligible

NVIDIA Corporation • Toronto, Canada

Full-time

A leading technology company in Toronto is seeking a Senior Software Engineer to architect and implement AI inference systems. The role involves optimizing GPU kernels and compilers, driving industr...Show more

Last updated: 24 days ago • Promoted

Staff AI DevOps Engineer – Cloud, MLOps & Kubernetes

Thomson Reuters • Toronto, Canada

Full-time

A leading global information provider is seeking a Staff Software Engineer – AI (DevOps) in Toronto.The ideal candidate will work on architecting and implementing AI-driven solutions and designing ...Show more

Last updated: 22 days ago • Promoted