Talent.com
Deployment Engineer, AI Inference
Deployment Engineer, AI InferenceCerebras Systems Inc. • Toronto, Canada
Deployment Engineer, AI Inference

Deployment Engineer, AI Inference

Cerebras Systems Inc. • Toronto, Canada
24 days ago
Job type
  • Full-time
Job description

About Cerebras Systems

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer‑scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of one device. This enables industry‑leading training and inference speeds and lets machine learning users run large‑scale ML applications without the hassle of managing hundreds of GPUs or TPUs.

Our customers include global corporations, national labs and top‑tier healthcare systems. In 2024 we launched Cerebras Inference, the fastest generative AI inference solution, over 10 times faster than GPU‑based hyperscale cloud inference services.

About The Role

We are seeking a highly skilled Deployment Engineer to build and operate cutting‑edge inference clusters on the world’s largest computer chip, the Wafer‑Scale Engine (WSE). You will play a critical role in ensuring reliable, efficient, and scalable deployment of AI inference workloads across our global infrastructure. On the operational side, you’ll own the rollout of new software versions, AI replica updates, and capacity reallocations across our custom‑built, high‑capacity datacenters. Beyond operations, you’ll drive improvements to telemetry, observability, and fully automated pipelines using advanced allocation strategies to maximize utilization of large‑scale computer fleets.

The ideal candidate combines hands‑on operation rigor with strong systems engineering skills and thrives on building resilient pipelines that keep pace with cutting‑edge AI models.

This role does not require 24 / 7 hour on‑call rotations.

Responsibilities

Deploy AI inference replicas and cluster software across multiple datacenters

Maximize capacity allocation and optimize replica placement using constraint‑solver algorithms

Operate bare‑metal inference infrastructure while supporting transition to K8S‑based platform

Develop and extend telemetry, observability and alerting solutions to ensure deployment reliability at scale

Develop and extend a fully automated deployment pipeline to support fast software updates and capacity reallocation at scale

Translate technical and customer needs into actionable requirements for the Dev Infra, Cluster, Platform and Core teams

Stay up to date with the latest advancements in AI compute infrastructure and related technologies

Skills And Requirements

2–5 years of experience operating on‑prem compute infrastructure (ideally in Machine Learning or High‑Performance Compute) or developing and managing complex AWS‑based infrastructure for hybrid deployments

Strong proficiency in Python for automation, orchestration, and deployment tooling

Solid understanding of Linux‑based systems and command‑line tools

Extensive knowledge of Docker containers and container orchestration platforms like K8S

Familiarity with spine‑leaf (Clos) networking architecture

Proficiency with telemetry and observability stacks such as Prometheus, InfluxDB and Grafana

Strong ownership mindset and accountability for complex deployments

Ability to work effectively in a fast‑paced environment

Location

SF Bay Area

Toronto

Why Join Cerebras

Build a breakthrough AI platform beyond the constraints of the GPU

Publish and open‑source cutting‑edge AI research

Work on one of the fastest AI supercomputers in the world

Enjoy job stability with startup vitality

Our simple, non‑corporate work culture respects individual beliefs

Apply Today and Become Part of the Forefront of Groundbreaking Advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. Inclusive teams build better products and companies, and we empower people to do their best work through continuous learning, growth, and support of those around them.

#J-18808-Ljbffr

Create a job alert for this search

Deployment Engineer AI Inference • Toronto, Canada

Similar jobs
AI Engineer

AI Engineer

TheAppLabb • Toronto, ON, Canada
Full-time
AI-powered digital solutions, mobile app development, and emerging technologies.We leverage data-driven insights to enhance digital experiences and drive business growth. We are looking for a Machin...Show more
Last updated: 30+ days ago • Promoted
Remote AI / ML Solutions Architect — AWS & GenAI Expert

Remote AI / ML Solutions Architect — AWS & GenAI Expert

Avahi • Toronto C6A, ON, Canada
Remote
Full-time
A cloud-first consulting company is looking for an experienced AI / ML Solutions Architect to join their remote-first team. The successful candidate will leverage AWS services to design scalable AI / ML...Show more
Last updated: 30+ days ago • Promoted
Senior LLMOps Engineer — Cloud AI Inference + Equity

Senior LLMOps Engineer — Cloud AI Inference + Equity

TEEMA Solutions Group • Toronto, Canada
Full-time
A rapid-growth technology firm in Toronto is seeking a Staff LLMOps Engineer to lead the design and optimization of large language model infrastructure on the cloud. The ideal candidate has over 6 y...Show more
Last updated: 16 days ago • Promoted
Strategic Deployment Lead for Enterprise AI (Toronto)

Strategic Deployment Lead for Enterprise AI (Toronto)

Peregrine • Toronto
Full-time
A technology solutions provider in Toronto seeks a Deployment Strategist to lead the product deployment and enhance customer impact. Ideal candidates will have 2–5 years of experience in deploying e...Show more
Last updated: 17 days ago • Promoted
Lead AI Engineer

Lead AI Engineer

Harnham • Toronto, ON, Canada
Full-time
Toronto, ON - 3 days onsite / week.CAD + bonus + LTI; 300,000 - 400,000 CAD total.Harnham is partnering with one of the most well known financial services companies, which is looking for an experienc...Show more
Last updated: 6 days ago • Promoted
Forward Deployed Engineer, Agentic AI Platform

Forward Deployed Engineer, Agentic AI Platform

Cohere • Toronto, Canada
Full-time
A technology company based in Toronto is looking for an Entry-Level Forward Deployed Engineer to help develop its AI workspace platform, North. This unique role involves building features and deploy...Show more
Last updated: 24 days ago • Promoted
AI Inference Digital Design Engineer

AI Inference Digital Design Engineer

Taalas • Toronto, Canada
Full-time
A tech company is seeking a passionate Digital Design Engineer in Toronto to work on complex AI Inference challenges.Candidates should have a degree in Electrical or Computer Engineering and profic...Show more
Last updated: 30+ days ago • Promoted
AI Engineer

AI Engineer

GEI Consultants • Markham, ON, Canada
Full-time
The AI Engineer is responsible for the development of AI solutions, typically leveraging pretrained models and copilots, to support GEI’s priority digital and AI initiatives.This role focuses...Show more
Last updated: 30+ days ago • Promoted
Cloud AI Inference Engineer - Golang

Cloud AI Inference Engineer - Golang

Huawei Canada • Markham
Full-time +1
A technology firm in York Region, Canada is seeking an Engineer for a 12-month contract.The role includes collaborating with senior engineers to design AI inference frameworks using Golang, optimiz...Show more
Last updated: 11 hours ago • Promoted • New!
AI Platform Engineer : Build & Evaluate AI Systems

AI Platform Engineer : Build & Evaluate AI Systems

Armilla AI • Toronto, Canada
Full-time
A leading AI-focused startup in Toronto is seeking an experienced AI Engineer to shape the future of AI risk management.In this pivotal role, you will be responsible for building core AI tools and ...Show more
Last updated: 30+ days ago • Promoted
Senior AI Platform Engineer — Remote

Senior AI Platform Engineer — Remote

Dayforce US, Inc. • Toronto, Canada
Remote
Full-time
A global HCM company is seeking an AI Platform Developer Senior to build scalable AI application platforms.This role focuses on designing infrastructure and developing tools to enhance developer wo...Show more
Last updated: 22 days ago • Promoted
Sr. Deployment Engineer, AI Inference

Sr. Deployment Engineer, AI Inference

Cerebras • Toronto
Full-time
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programm...Show more
Last updated: 11 hours ago • Promoted • New!
Forward-Deployed AI Engineer for Enterprise Solutions

Forward-Deployed AI Engineer for Enterprise Solutions

Google Inc. • Toronto, Canada
Full-time
A leading technology company is seeking a Forward Deployed Engineer to utilize Generative AI in solving customer challenges. The role involves software development, deployment, and collaboration wit...Show more
Last updated: 19 days ago • Promoted
AI Systems & Inference Frameworks Engineer

AI Systems & Inference Frameworks Engineer

adaption • Toronto, Canada
Full-time
About Us Most AI is frozen in place - it doesn't adapt to the world.Our mandate is to build efficient intelligence that evolves in real‑time. Our vision is AI systems that are flexible, personalized...Show more
Last updated: 22 days ago • Promoted
Remote-First LLM AI Research Engineer (NLP Innovator)

Remote-First LLM AI Research Engineer (NLP Innovator)

Upfeat Media, Inc. • Toronto, Canada
Remote
Full-time
A leading company in applied AI seeks a Large Language Model AI Research Engineer passionate about NLP and innovative solutions. This role involves developing datasets, engaging with cutting-edge te...Show more
Last updated: 24 days ago • Promoted
AI Solutions Engineer - Lead Deployment & Impact

AI Solutions Engineer - Lead Deployment & Impact

Inbenta • Toronto, Canada
Full-time
A global technology company in Toronto is seeking an AI Solutions Engineer to lead technical engagement across the customer lifecycle. The role involves solution design, implementation, and ongoing ...Show more
Last updated: 22 days ago • Promoted
Senior AI Inference Systems Engineer — Equity Eligible

Senior AI Inference Systems Engineer — Equity Eligible

NVIDIA Corporation • Toronto, Canada
Full-time
A leading technology company in Toronto is seeking a Senior Software Engineer to architect and implement AI inference systems. The role involves optimizing GPU kernels and compilers, driving industr...Show more
Last updated: 24 days ago • Promoted
Staff AI DevOps Engineer – Cloud, MLOps & Kubernetes

Staff AI DevOps Engineer – Cloud, MLOps & Kubernetes

Thomson Reuters • Toronto, Canada
Full-time
A leading global information provider is seeking a Staff Software Engineer – AI (DevOps) in Toronto.The ideal candidate will work on architecting and implementing AI-driven solutions and designing ...Show more
Last updated: 22 days ago • Promoted