Deployment Engineer, AI InferenceCerebras Systems Inc. • Toronto, Canada

Deployment Engineer, AI Inference

Cerebras Systems Inc. • Toronto, Canada

Il y a 25 jours

Type de contrat

Temps plein

Description de poste

About Cerebras Systems

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer‑scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of one device. This enables industry‑leading training and inference speeds and lets machine learning users run large‑scale ML applications without the hassle of managing hundreds of GPUs or TPUs.

Our customers include global corporations, national labs and top‑tier healthcare systems. In 2024 we launched Cerebras Inference, the fastest generative AI inference solution, over 10 times faster than GPU‑based hyperscale cloud inference services.

About The Role

We are seeking a highly skilled Deployment Engineer to build and operate cutting‑edge inference clusters on the world’s largest computer chip, the Wafer‑Scale Engine (WSE). You will play a critical role in ensuring reliable, efficient, and scalable deployment of AI inference workloads across our global infrastructure. On the operational side, you’ll own the rollout of new software versions, AI replica updates, and capacity reallocations across our custom‑built, high‑capacity datacenters. Beyond operations, you’ll drive improvements to telemetry, observability, and fully automated pipelines using advanced allocation strategies to maximize utilization of large‑scale computer fleets.

The ideal candidate combines hands‑on operation rigor with strong systems engineering skills and thrives on building resilient pipelines that keep pace with cutting‑edge AI models.

This role does not require 24 / 7 hour on‑call rotations.

Responsibilities

Deploy AI inference replicas and cluster software across multiple datacenters

Maximize capacity allocation and optimize replica placement using constraint‑solver algorithms

Operate bare‑metal inference infrastructure while supporting transition to K8S‑based platform

Develop and extend telemetry, observability and alerting solutions to ensure deployment reliability at scale

Develop and extend a fully automated deployment pipeline to support fast software updates and capacity reallocation at scale

Translate technical and customer needs into actionable requirements for the Dev Infra, Cluster, Platform and Core teams

Stay up to date with the latest advancements in AI compute infrastructure and related technologies

Skills And Requirements

2–5 years of experience operating on‑prem compute infrastructure (ideally in Machine Learning or High‑Performance Compute) or developing and managing complex AWS‑based infrastructure for hybrid deployments

Strong proficiency in Python for automation, orchestration, and deployment tooling

Solid understanding of Linux‑based systems and command‑line tools

Extensive knowledge of Docker containers and container orchestration platforms like K8S

Familiarity with spine‑leaf (Clos) networking architecture

Proficiency with telemetry and observability stacks such as Prometheus, InfluxDB and Grafana

Strong ownership mindset and accountability for complex deployments

Ability to work effectively in a fast‑paced environment

Location

SF Bay Area

Toronto

Why Join Cerebras

Build a breakthrough AI platform beyond the constraints of the GPU

Publish and open‑source cutting‑edge AI research

Work on one of the fastest AI supercomputers in the world

Enjoy job stability with startup vitality

Our simple, non‑corporate work culture respects individual beliefs

Apply Today and Become Part of the Forefront of Groundbreaking Advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. Inclusive teams build better products and companies, and we empower people to do their best work through continuous learning, growth, and support of those around them.

#J-18808-Ljbffr

Créer une alerte emploi pour cette recherche

Deployment Engineer AI Inference • Toronto, Canada

Offres similaires

Remote AI / ML Solutions Architect — AWS & GenAI Expert

Avahi • Toronto C6A, ON, Canada

Télétravail

Temps plein

A cloud-first consulting company is looking for an experienced AI / ML Solutions Architect to join their remote-first team. The successful candidate will leverage AWS services to design scalable AI / ML...Voir plus

Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée

Senior LLMOps Engineer — Cloud AI Inference + Equity

TEEMA Solutions Group • Toronto, Canada

Temps plein

A rapid-growth technology firm in Toronto is seeking a Staff LLMOps Engineer to lead the design and optimization of large language model infrastructure on the cloud. The ideal candidate has over 6 y...Voir plus

Dernière mise à jour : il y a 17 jours • Offre sponsorisée

Lead AI Engineer

Harnham • Toronto, ON, Canada

Temps plein

Toronto, ON - 3 days onsite / week.CAD + bonus + LTI; 300,000 - 400,000 CAD total.Harnham is partnering with one of the most well known financial services companies, which is looking for an experienc...Voir plus

Dernière mise à jour : il y a 6 jours • Offre sponsorisée

Principal Staff Engineer – AI Infrastructure - AI / ML Leader

Andiamo • Toronto C6A, ON, Canada

Temps plein +1

Principal Staff Engineer - AI Infrastructure.This role sits at the intersection of large-scale distributed systems and cutting-edge machine learning, powering the platforms that enable researchers ...Voir plus

Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée

Forward Deployed Engineer, Agentic AI Platform

Cohere • Toronto, Canada

Temps plein

A technology company based in Toronto is looking for an Entry-Level Forward Deployed Engineer to help develop its AI workspace platform, North. This unique role involves building features and deploy...Voir plus

Dernière mise à jour : il y a 25 jours • Offre sponsorisée

AI Inference Digital Design Engineer

Taalas • Toronto, Canada

Temps plein

A tech company is seeking a passionate Digital Design Engineer in Toronto to work on complex AI Inference challenges.Candidates should have a degree in Electrical or Computer Engineering and profic...Voir plus

Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée

Lead AI Research Engineer, NLP / ML — Hybrid

Refinitiv • Toronto, Canada

Temps plein

A leading technology firm in Toronto seeks a Lead Research Engineer to provide technical leadership and develop innovative AI solutions. You will lead cross-functional teams, work with cutting-edge ...Voir plus

Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée

Azure AI Architect — GenAI & MLOps Leader

TD • Toronto C6A, ON, Canada

Temps plein

A leading financial institution is seeking an experienced AI Architect to oversee the design and implementation of enterprise-grade AI solutions using Microsoft Azure in Toronto.In this critical ro...Voir plus

Dernière mise à jour : il y a 23 jours • Offre sponsorisée

AI Platform Engineer : Build & Evaluate AI Systems

Armilla AI • Toronto, Canada

Temps plein

A leading AI-focused startup in Toronto is seeking an experienced AI Engineer to shape the future of AI risk management.In this pivotal role, you will be responsible for building core AI tools and ...Voir plus

Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée

Senior AI Platform Engineer — Remote

Dayforce US, Inc. • Toronto, Canada

Télétravail

Temps plein

A global HCM company is seeking an AI Platform Developer Senior to build scalable AI application platforms.This role focuses on designing infrastructure and developing tools to enhance developer wo...Voir plus

Dernière mise à jour : il y a 22 jours • Offre sponsorisée

Sr. Deployment Engineer, AI Inference

Cerebras • Toronto

Temps plein

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programm...Voir plus

Dernière mise à jour : il y a 19 heures • Offre sponsorisée • Nouvelle offre

Forward-Deployed AI Engineer for Enterprise Solutions

Google Inc. • Toronto, Canada

Temps plein

A leading technology company is seeking a Forward Deployed Engineer to utilize Generative AI in solving customer challenges. The role involves software development, deployment, and collaboration wit...Voir plus

Dernière mise à jour : il y a 19 jours • Offre sponsorisée

AI / ML Platform Engineer – Intelligence Accelerator

Okta • Toronto C6A, ON, Canada

Temps plein

A leading identity management company is seeking an AI / ML Engineer II to join their Intelligence Accelerator team in Toronto. This role will involve designing scalable ML infrastructure and collabor...Voir plus

Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée

AI Systems & Inference Frameworks Engineer

adaption • Toronto, Canada

Temps plein

About Us Most AI is frozen in place - it doesn't adapt to the world.Our mandate is to build efficient intelligence that evolves in real‑time. Our vision is AI systems that are flexible, personalized...Voir plus

Dernière mise à jour : il y a 23 jours • Offre sponsorisée

Remote-First LLM AI Research Engineer (NLP Innovator)

Upfeat Media, Inc. • Toronto, Canada

Télétravail

Temps plein

A leading company in applied AI seeks a Large Language Model AI Research Engineer passionate about NLP and innovative solutions. This role involves developing datasets, engaging with cutting-edge te...Voir plus

Dernière mise à jour : il y a 25 jours • Offre sponsorisée

AI Solutions Engineer - Lead Deployment & Impact

Inbenta • Toronto, Canada

Temps plein

A global technology company in Toronto is seeking an AI Solutions Engineer to lead technical engagement across the customer lifecycle. The role involves solution design, implementation, and ongoing ...Voir plus

Dernière mise à jour : il y a 23 jours • Offre sponsorisée

Senior AI Inference Systems Engineer — Equity Eligible

NVIDIA Corporation • Toronto, Canada

Temps plein

A leading technology company in Toronto is seeking a Senior Software Engineer to architect and implement AI inference systems. The role involves optimizing GPU kernels and compilers, driving industr...Voir plus

Dernière mise à jour : il y a 25 jours • Offre sponsorisée

Staff AI DevOps Engineer – Cloud, MLOps & Kubernetes

Thomson Reuters • Toronto, Canada

Temps plein

A leading global information provider is seeking a Staff Software Engineer – AI (DevOps) in Toronto.The ideal candidate will work on architecting and implementing AI-driven solutions and designing ...Voir plus

Dernière mise à jour : il y a 23 jours • Offre sponsorisée