Talent.com
Boson AI
Site Reliability Engineer, AI/ML InfrastructureBoson AI • Toronto, ON, CA
Site Reliability Engineer, AI/ML Infrastructure

Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto, ON, CA
30+ days ago
Job type
  • Full-time
Job description

Overview

We2;re looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters aroundour Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers.

Youll be hands-on with the full lifecycle of HPC infrastructure: planning, building, testing, deploying, and keeping everything running smoothly. That means troubleshooting issues as they arise, monitoring performance, developing automation to make our lives easier, and working closely with engineering and science teams to ensure they have what they need. Youll also help us plan for future capacity and evaluate new technologies as we continue to scale.

Responsibilities

  • Manage and optimize HPC cluster operations
  • Deploy and maintain infrastructure-as-code solutions
  • Support ML/research teams with cluster usage optimization
  • Operate, troubleshoot and optimize Ceph storage clusters
  • Develop automation and tooling

Minimum Qualifications

  • 5+ years of experience in SRE or HPC operations
  • Proficiency in Linux systems administration (Ubuntu/Debian)
  • Experience with Kubernetes and container orchestration
  • Experience with Ceph >1PB deployments and maintenance
  • Knowledge of security best practices in multi-tenant environments
  • Understanding of L2/L3 networking fundamentals
  • Skilled in Python and Bash scripting

Preferred Qualifications

  • Experience with infrastructure-as-code tools (Ansible/Terraform)
  • Experience with GitOps (Helm, ArgoCD)
  • Strong grasp of RDMA, InfiniBand, and GPUDirect technologies
  • Familiarity with deep learning frameworks such as PyTorch and TensorFlow
  • Familiarity in at least one cloud platform: AWS, Azure or GCP

If youre a natural problem-solver with a passion for continuous learning, wed love to hear from you.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

#J-18808-Ljbffr
Create a job alert for this search

Site Reliability Engineer, AI/ML Infrastructure • Toronto, ON, CA

Similar jobs

Senior Site Reliability Engineer

Guidewire Softwaretoronto, on, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Senior Site Reliability Engineer, Kong Konnect

Kong Inc.toronto, on, Canada
Full-time

Senior Site Reliability Engineer, Kong Konnect.This range is provided by Kong Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Are you ready ... Show more

 • Promoted

Generative AI Engineer — Azure & LLM Solutions

GEI Consultantsmarkham, york region, Canada
Full-time

A leading engineering and consulting firm in Ontario seeks an AI Engineer to develop and deploy AI solutions.The ideal candidate will have extensive software engineering experience and skills in Ge... Show more

 • Promoted

Site Reliability Engineer

Socket.devtoronto, on, Canada
Full-time

We are seeking a Senior Consultant in Site Reliability Engineering (Network SRE) to lead network-centric reliability practices across the Shared Platform ecosystem.This role focuses on ensuring res... Show more

 • Promoted

Remote Aerospace AI Engineer — Train & Improve Models

DataAnnotationmarkham, york region, Canada
Remote
Full-time

An innovative company is seeking a skilled Aerospace Engineer to enhance AI models through the application of physics.This role allows you to work remotely, choose your projects, and set your own s... Show more

 • Promoted

Site Reliability Engineer Role at Magnet Forensics

Magnet ForensicsToronto, ON, CA
Full-time

Take your expertise to the next level as a Senior Site Reliability Engineer with Magnet Forensics.This role involves hands-on AWS and Kubernetes management to uphold our SaaS platform’s reliability... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, ON, CA
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office].The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Site Reliability Engineer

Future Secure AIToronto, ON, CA
Full-time

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Show more

 • Promoted

Lead Site Reliability Engineer at iManage

iManageToronto, ON, CA
Full-time

Advance your career as a Lead Site Reliability Engineer at iManage, focused on maintaining and enhancing cloud resilience while enjoying flexible work arrangements.You will play a crucial role in d... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

AI Systems Engineer – AI Model (Training & Inference)

AMDmarkham, on, Canada
Full-time

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems.Grounded in a culture of innovatio... Show more

 • Promoted

Senior Engineer-Cloud AI Infrastructure

Huawei Canadamarkham, york region, Canada
Permanent

Huawei Canada has an immediate permanent opening for a Senior Engineer.Established in 2014, the Distributed Scheduling and Data Engine Lab is Huawei Cloud's technical innovation center in Canada.Th... Show more

 • Promoted

Site Reliability Engineer

Momentum Financial Services GroupToronto, ON, CA
Full-time

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Show more

 • Promoted

Site Reliability Engineer for Cloud Infrastructure Management

NewtonToronto, ON, CA
Full-time

Be a pivotal Site Reliability Engineer focused on improving infrastructure resilience and reliability.Collaborate remotely to drive operational success and enhance system performance in a dynamic e... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsToronto, ON, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more