Talent.com
Cohere
Senior ML Systems Engineer, Frameworks & ToolingCohere • Montreal, Montreal (administrative region), CA
Senior ML Systems Engineer, Frameworks & Tooling

Senior ML Systems Engineer, Frameworks & Tooling

Cohere • Montreal, Montreal (administrative region), CA
30+ days ago
Job type
  • Full-time
Job description

Join to apply for the Senior ML Systems Engineer, Frameworks & Tooling role at Cohere

Who are we?

Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. We like to work hard and move fast to do what’s best for our customers.

Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each person is one of the best in the world at what they do. We believe that a diverse range of perspectives is a requirement for building great products.

Join us on our mission and shape the future!

We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier‑scale language models. This role sits at the intersection of large‑scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training — and build the tooling that connects research ideas to thousands of GPUs.

If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact.

What You’ll Work On

  • Build and own the training framework responsible for large‑scale LLM training.
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing).
  • Improve training throughput and stability on multi‑node clusters (e.g., GB200/300, AMD, H200/100).
  • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
  • Collaborate closely with infra teams to ensure Slurm setups, container environments, and hardware configurations support high‑performance training.
  • Investigate and resolve performance bottlenecks across the ML systems stack.
  • Build robust systems that ensure reproducible, debuggable, large‑scale runs.

You Might Be a Good Fit If You Have

  • Strong engineering experience in large‑scale distributed training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
  • Experience with multi‑node cluster orchestration (Slurm, Ray, Kubernetes, or similar).
  • Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.
  • Experience working with containerized environments (Docker, Singularity/Apptainer).
  • A track record of building tools that increase developer velocity for ML teams.
  • Excellent judgment around trade‑offs: performance versus complexity, research velocity versus maintainability.
  • Strong collaboration skills — you’ll work closely with infra, research, and deployment teams.

Nice to Have

  • Experience with training LLMs or other large transformer architectures.
  • Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).
  • Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).
  • Experience with data pipeline optimization, sharded datasets, or caching strategies.
  • Background in performance engineering, profiling, or low‑level systems.

Bonus: paper at top‑tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP).

Why Join Us

  • You’ll work on some of the most challenging and consequential ML systems problems today.
  • You’ll collaborate with a world‑class team working fast and at scale.
  • You’ll have end‑to‑end ownership over critical components of the training stack.
  • You’ll shape the next generation of infrastructure for frontier‑scale models.
  • You’ll build tools and systems that directly accelerate research and model quality.

Sample Projects

  • Build a high‑performance data loading and caching pipeline.
  • Implement performance profiling across the ML systems stack
  • Develop internal metrics and monitoring for training runs.
  • Build reproducibility and regression testing infrastructure.
  • Develop a performant fault‑tolerant distributed checkpointing system.

If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!

We value and celebrate diversity and strive to create an inclusive work environment for all. We welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.

Full‑Time Employees At Cohere Enjoy These Perks

  • 🤝 An open and inclusive culture and work environment
  • 🧑💻 Work closely with a team on the cutting edge of AI research
  • 🍽 Weekly lunch stipend, in‑office lunches & snacks
  • 🦷 Full health and dental benefits, including a separate budget to take care of your mental health
  • 🐣 100 % Parental Leave top‑up for up to 6 months
  • 🎨 Personal enrichment benefits towards arts and culture, fitness and well‑being, quality time, and workspace improvement
  • 🏙 Remote‑flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co‑working stipend
  • ✈️ 6 weeks of vacation (30 working days!)

Seniority level

  • Mid‑Senior level

Employment type

  • Full‑time

Job function

  • Information Technology

Industries

  • Software Development

Location: Montreal, Quebec, Canada

#J-18808-Ljbffr
Create a job alert for this search

Senior ML Systems Engineer, Frameworks & Tooling • Montreal, Montreal (administrative region), CA

Similar jobs

MuleSoft Platform Support Engineer

Quantum World Technologies Inc.Montreal (administrative region), QC, CA
Full-time

Driving Talent Acquisition Across Canada || 26K+ Trusted Connections || Delivering Top IT Talent for Tomorrow’s Innovations.Role: Lead MuleSoft Platform Support Engineer.Provide platform support fo... Show more

 • Promoted

Scale ML Infra & Distributed Systems Engineer

Hunter BondMontréal, Montreal (administrative region), Canada
Full-time

A leading tech firm in Montreal is seeking exceptional Software Engineers to join their world-class engineering team focused on complex distributed systems and ML infrastructure.Ideal candidates wi... Show more

 • Promoted

Senior Ml Engineer Role At Surveymonkey

SurveyMonkeyRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Elevate your career as a Senior Software Engineer II at SurveyMonkey, specializing in creating scalable ML pipelines.Work hybrid in Ottawa with a focus on Python and AWS technologies.In this pivota... Show more

 • Promoted

Senior Manufacturing Systems Engineer

L3Harris Technologiesmontreal (administrative region), qc, Canada
Full-time

Senior Manufacturing Systems Engineer.Between $100,500 - $150,500 CDN annually.The Operations Engineering Systems Engineering Lead (OSE) serves as the technical lead for production systems, bridgin... Show more

 • Promoted

Engineering Manager, MAAS — Lead Distributed Systems

CanonicalMontreal (administrative region), QC, CA
Full-time

A leading open source software provider is seeking an Engineering Manager to lead the MAAS team.The successful candidate will have a solid background in software development using Python or Golang ... Show more

 • Promoted

Senior Software Engineer, Backend for ML

Coveo Solutions Inc.Montreal (administrative region), QC, CA
Full-time

Transform machine learning practices at Coveo as a Senior Software Engineer.Build infrastructure that bridges experimentation and production while optimizing ML workflows and developer experiences.... Show more

 • Promoted

Remote ML Architect - AWS Cloud Native

CaylentMontreal (administrative region), QC, CA
Remote
Full-time

A cloud-native services company is seeking a Machine Learning Architect to join its Cloud Native Applications team.This role requires expertise in ML system design and AWS solutions.You will work c... Show more

 • Promoted

Principal MLOPs Engineer (Canada)

Rackspace TechnologyMontreal (administrative region), QC, CA
Full-time

Be among the first 25 applicants.Get AI-powered advice on this job and more exclusive features.We are looking for a seasoned Principal ML OPS Engineer to architect, build, and optimize ML inference... Show more

 • Promoted

Bilingual Project Engineer - Medium Systems

Camfil Power SystemsLaval (administrative region), QC, CA
Full-time

A leader in air filtration technology in Laval is seeking an engineering professional to manage project compliance and team coordination.The ideal candidate will have 3-5 years of experience in eng... Show more

 • Promoted

Senior Software Engineer - High-Throughput ML Personalization

AC780Montréal, Montreal (administrative region), Canada
Full-time

A rapidly growing company is seeking a Software Engineer to enhance its personalization platform.This role involves designing complex systems and collaborating with machine learning engineers to de... Show more

 • Promoted

Senior ML/DL Developer

Stay22Montreal (administrative region), QC, CA
Full-time

Neuro fournit les fondations techniques qui alimentent nos moteurs principaux, notamment.Neuro Squad, vous serez responsable de concevoir et d’architecturer l’intelligence derrière ces produits.Ce ... Show more

 • Promoted

Senior Distributed Diffusion ML Engineer — Remote

Bagel LabsMontreal (administrative region), QC, CA
Remote
Full-time

A leading machine learning research lab is looking for an expert in distributed systems to design scalable infrastructure for diffusion models.The ideal candidate will have substantial experience i... Show more

 • Promoted

Ingénieur ML et données Senior chez Mistplay

MistplayMontreal (administrative region), QC, CA
Full-time

Rejoignez Mistplay en tant qu'Ingénieur ML et données Senior dans un rôle hybride à Toronto ou Montréal.Votre mission sera de mettre en place une plateforme performante favorisant l'engagement des ... Show more

 • Promoted

ML Engineer (Speech-to-Speech) — Subject Matter Expert

VosynMontreal (administrative region), QC, CA
Full-time

ML Engineer (Speech-to-Speech) — Subject Matter Expert.At Vosyn, we embrace the exciting, game-changing world of Artificial Intelligence, driving innovation and pioneering impactful projects across... Show more

 • Promoted

Intermediate Enterprise and Systems Architect/Modeler / Architecte/Modélisateur(trice) interméd[...]

The Weir Group PLCMontreal
Full-time +1

Intermediate Enterprise and Systems Architect/Modeler - NETE.Location: LaSalle, Ottawa, Gatineau and Halifax.Job Type: Permanent Full-time, Onsite work.The position of Intermediate Enterprise and S... Show more

 • Promoted

Digital-Analog Systems Lead Engineer Position

SyntronicMontreal
Full-time

Innovate with Syntronic as a Digital-Analog Systems Lead Engineer.Your expertise will drive projects in aerospace and telecommunications design, shaping the future of technology on a global scale.I... Show more

 • Promoted

Senior, Ml Engineer - Ml Ops Framework - C$141,500 - C$212,300 A Year

TorcWestmount, Canada
Full-time

Seeking a Senior ML Engineer to build and maintain ML frameworks for autonomous vehicle solutions, focusing on cloud tooling, data operations, and ML development pipelines using AWS and Python. Show more

 • Promoted

Giesecke+Devrient Senior AI Systems Engineer

Giesecke+Devrientmontreal (administrative region), qc, Canada
Full-time

Drive Generative AI innovation at Giesecke+Devrient as a Senior Applied AI Engineer in the newly established AI Hub.Combine your strong Python skills with advanced AI engineering knowledge.This key... Show more

 • Promoted

Software Engineer for ML Inference

BasetenMontreal (administrative region), QC, CA
Full-time

Step into innovation as a Software Engineer focused on ML Inference at Baseten, where your skills will transform AI applications.You’ll work on high-impact components using Python and Go to optimiz... Show more

 • Promoted

Senior ML Engineer - Graph & Transformer Systems

SAPMontreal
Full-time

A leading software company is looking for a Senior Machine Learning Engineer in Montreal to lead the development of scalable ML systems.The role involves mentoring engineers and optimizing data pip... Show more