Talent.com
Deployment Engineer, AI Inference
Deployment Engineer, AI InferenceCerebras Systems Inc. • Toronto, Canada
Deployment Engineer, AI Inference

Deployment Engineer, AI Inference

Cerebras Systems Inc. • Toronto, Canada
Il y a 25 jours
Type de contrat
  • Temps plein
Description de poste

About Cerebras Systems

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer‑scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of one device. This enables industry‑leading training and inference speeds and lets machine learning users run large‑scale ML applications without the hassle of managing hundreds of GPUs or TPUs.

Our customers include global corporations, national labs and top‑tier healthcare systems. In 2024 we launched Cerebras Inference, the fastest generative AI inference solution, over 10 times faster than GPU‑based hyperscale cloud inference services.

About The Role

We are seeking a highly skilled Deployment Engineer to build and operate cutting‑edge inference clusters on the world’s largest computer chip, the Wafer‑Scale Engine (WSE). You will play a critical role in ensuring reliable, efficient, and scalable deployment of AI inference workloads across our global infrastructure. On the operational side, you’ll own the rollout of new software versions, AI replica updates, and capacity reallocations across our custom‑built, high‑capacity datacenters. Beyond operations, you’ll drive improvements to telemetry, observability, and fully automated pipelines using advanced allocation strategies to maximize utilization of large‑scale computer fleets.

The ideal candidate combines hands‑on operation rigor with strong systems engineering skills and thrives on building resilient pipelines that keep pace with cutting‑edge AI models.

This role does not require 24 / 7 hour on‑call rotations.

Responsibilities

Deploy AI inference replicas and cluster software across multiple datacenters

Maximize capacity allocation and optimize replica placement using constraint‑solver algorithms

Operate bare‑metal inference infrastructure while supporting transition to K8S‑based platform

Develop and extend telemetry, observability and alerting solutions to ensure deployment reliability at scale

Develop and extend a fully automated deployment pipeline to support fast software updates and capacity reallocation at scale

Translate technical and customer needs into actionable requirements for the Dev Infra, Cluster, Platform and Core teams

Stay up to date with the latest advancements in AI compute infrastructure and related technologies

Skills And Requirements

2–5 years of experience operating on‑prem compute infrastructure (ideally in Machine Learning or High‑Performance Compute) or developing and managing complex AWS‑based infrastructure for hybrid deployments

Strong proficiency in Python for automation, orchestration, and deployment tooling

Solid understanding of Linux‑based systems and command‑line tools

Extensive knowledge of Docker containers and container orchestration platforms like K8S

Familiarity with spine‑leaf (Clos) networking architecture

Proficiency with telemetry and observability stacks such as Prometheus, InfluxDB and Grafana

Strong ownership mindset and accountability for complex deployments

Ability to work effectively in a fast‑paced environment

Location

SF Bay Area

Toronto

Why Join Cerebras

Build a breakthrough AI platform beyond the constraints of the GPU

Publish and open‑source cutting‑edge AI research

Work on one of the fastest AI supercomputers in the world

Enjoy job stability with startup vitality

Our simple, non‑corporate work culture respects individual beliefs

Apply Today and Become Part of the Forefront of Groundbreaking Advancements in AI!

Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. Inclusive teams build better products and companies, and we empower people to do their best work through continuous learning, growth, and support of those around them.

#J-18808-Ljbffr

Créer une alerte emploi pour cette recherche

Deployment Engineer AI Inference • Toronto, Canada

Offres similaires
Remote AI / ML Solutions Architect — AWS & GenAI Expert

Remote AI / ML Solutions Architect — AWS & GenAI Expert

Avahi • Toronto C6A, ON, Canada
Télétravail
Temps plein
A cloud-first consulting company is looking for an experienced AI / ML Solutions Architect to join their remote-first team. The successful candidate will leverage AWS services to design scalable AI / ML...Voir plus
Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée
Senior LLMOps Engineer — Cloud AI Inference + Equity

Senior LLMOps Engineer — Cloud AI Inference + Equity

TEEMA Solutions Group • Toronto, Canada
Temps plein
A rapid-growth technology firm in Toronto is seeking a Staff LLMOps Engineer to lead the design and optimization of large language model infrastructure on the cloud. The ideal candidate has over 6 y...Voir plus
Dernière mise à jour : il y a 17 jours • Offre sponsorisée
Lead AI Engineer

Lead AI Engineer

Harnham • Toronto, ON, Canada
Temps plein
Toronto, ON - 3 days onsite / week.CAD + bonus + LTI; 300,000 - 400,000 CAD total.Harnham is partnering with one of the most well known financial services companies, which is looking for an experienc...Voir plus
Dernière mise à jour : il y a 6 jours • Offre sponsorisée
Principal Staff Engineer – AI Infrastructure - AI / ML Leader

Principal Staff Engineer – AI Infrastructure - AI / ML Leader

Andiamo • Toronto C6A, ON, Canada
Temps plein +1
Principal Staff Engineer - AI Infrastructure.This role sits at the intersection of large-scale distributed systems and cutting-edge machine learning, powering the platforms that enable researchers ...Voir plus
Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée
Forward Deployed Engineer, Agentic AI Platform

Forward Deployed Engineer, Agentic AI Platform

Cohere • Toronto, Canada
Temps plein
A technology company based in Toronto is looking for an Entry-Level Forward Deployed Engineer to help develop its AI workspace platform, North. This unique role involves building features and deploy...Voir plus
Dernière mise à jour : il y a 25 jours • Offre sponsorisée
AI Inference Digital Design Engineer

AI Inference Digital Design Engineer

Taalas • Toronto, Canada
Temps plein
A tech company is seeking a passionate Digital Design Engineer in Toronto to work on complex AI Inference challenges.Candidates should have a degree in Electrical or Computer Engineering and profic...Voir plus
Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée
Lead AI Research Engineer, NLP / ML — Hybrid

Lead AI Research Engineer, NLP / ML — Hybrid

Refinitiv • Toronto, Canada
Temps plein
A leading technology firm in Toronto seeks a Lead Research Engineer to provide technical leadership and develop innovative AI solutions. You will lead cross-functional teams, work with cutting-edge ...Voir plus
Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée
Azure AI Architect — GenAI & MLOps Leader

Azure AI Architect — GenAI & MLOps Leader

TD • Toronto C6A, ON, Canada
Temps plein
A leading financial institution is seeking an experienced AI Architect to oversee the design and implementation of enterprise-grade AI solutions using Microsoft Azure in Toronto.In this critical ro...Voir plus
Dernière mise à jour : il y a 23 jours • Offre sponsorisée
AI Platform Engineer : Build & Evaluate AI Systems

AI Platform Engineer : Build & Evaluate AI Systems

Armilla AI • Toronto, Canada
Temps plein
A leading AI-focused startup in Toronto is seeking an experienced AI Engineer to shape the future of AI risk management.In this pivotal role, you will be responsible for building core AI tools and ...Voir plus
Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée
Senior AI Platform Engineer — Remote

Senior AI Platform Engineer — Remote

Dayforce US, Inc. • Toronto, Canada
Télétravail
Temps plein
A global HCM company is seeking an AI Platform Developer Senior to build scalable AI application platforms.This role focuses on designing infrastructure and developing tools to enhance developer wo...Voir plus
Dernière mise à jour : il y a 22 jours • Offre sponsorisée
Sr. Deployment Engineer, AI Inference

Sr. Deployment Engineer, AI Inference

Cerebras • Toronto
Temps plein
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs.Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programm...Voir plus
Dernière mise à jour : il y a 19 heures • Offre sponsorisée • Nouvelle offre
Forward-Deployed AI Engineer for Enterprise Solutions

Forward-Deployed AI Engineer for Enterprise Solutions

Google Inc. • Toronto, Canada
Temps plein
A leading technology company is seeking a Forward Deployed Engineer to utilize Generative AI in solving customer challenges. The role involves software development, deployment, and collaboration wit...Voir plus
Dernière mise à jour : il y a 19 jours • Offre sponsorisée
AI / ML Platform Engineer – Intelligence Accelerator

AI / ML Platform Engineer – Intelligence Accelerator

Okta • Toronto C6A, ON, Canada
Temps plein
A leading identity management company is seeking an AI / ML Engineer II to join their Intelligence Accelerator team in Toronto. This role will involve designing scalable ML infrastructure and collabor...Voir plus
Dernière mise à jour : il y a plus de 30 jours • Offre sponsorisée
AI Systems & Inference Frameworks Engineer

AI Systems & Inference Frameworks Engineer

adaption • Toronto, Canada
Temps plein
About Us Most AI is frozen in place - it doesn't adapt to the world.Our mandate is to build efficient intelligence that evolves in real‑time. Our vision is AI systems that are flexible, personalized...Voir plus
Dernière mise à jour : il y a 23 jours • Offre sponsorisée
Remote-First LLM AI Research Engineer (NLP Innovator)

Remote-First LLM AI Research Engineer (NLP Innovator)

Upfeat Media, Inc. • Toronto, Canada
Télétravail
Temps plein
A leading company in applied AI seeks a Large Language Model AI Research Engineer passionate about NLP and innovative solutions. This role involves developing datasets, engaging with cutting-edge te...Voir plus
Dernière mise à jour : il y a 25 jours • Offre sponsorisée
AI Solutions Engineer - Lead Deployment & Impact

AI Solutions Engineer - Lead Deployment & Impact

Inbenta • Toronto, Canada
Temps plein
A global technology company in Toronto is seeking an AI Solutions Engineer to lead technical engagement across the customer lifecycle. The role involves solution design, implementation, and ongoing ...Voir plus
Dernière mise à jour : il y a 23 jours • Offre sponsorisée
Senior AI Inference Systems Engineer — Equity Eligible

Senior AI Inference Systems Engineer — Equity Eligible

NVIDIA Corporation • Toronto, Canada
Temps plein
A leading technology company in Toronto is seeking a Senior Software Engineer to architect and implement AI inference systems. The role involves optimizing GPU kernels and compilers, driving industr...Voir plus
Dernière mise à jour : il y a 25 jours • Offre sponsorisée
Staff AI DevOps Engineer – Cloud, MLOps & Kubernetes

Staff AI DevOps Engineer – Cloud, MLOps & Kubernetes

Thomson Reuters • Toronto, Canada
Temps plein
A leading global information provider is seeking a Staff Software Engineer – AI (DevOps) in Toronto.The ideal candidate will work on architecting and implementing AI-driven solutions and designing ...Voir plus
Dernière mise à jour : il y a 23 jours • Offre sponsorisée