Talent.com
Boson AI
Site Reliability Engineer, AI/ML InfrastructureBoson AI • Toronto, ON, CA
Site Reliability Engineer, AI/ML Infrastructure

Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto, ON, CA
30+ days ago
Job type
  • Full-time
Job description

Overview

We2;re looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters aroundour Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers.

Youll be hands-on with the full lifecycle of HPC infrastructure: planning, building, testing, deploying, and keeping everything running smoothly. That means troubleshooting issues as they arise, monitoring performance, developing automation to make our lives easier, and working closely with engineering and science teams to ensure they have what they need. Youll also help us plan for future capacity and evaluate new technologies as we continue to scale.

Responsibilities

  • Manage and optimize HPC cluster operations
  • Deploy and maintain infrastructure-as-code solutions
  • Support ML/research teams with cluster usage optimization
  • Operate, troubleshoot and optimize Ceph storage clusters
  • Develop automation and tooling

Minimum Qualifications

  • 5+ years of experience in SRE or HPC operations
  • Proficiency in Linux systems administration (Ubuntu/Debian)
  • Experience with Kubernetes and container orchestration
  • Experience with Ceph >1PB deployments and maintenance
  • Knowledge of security best practices in multi-tenant environments
  • Understanding of L2/L3 networking fundamentals
  • Skilled in Python and Bash scripting

Preferred Qualifications

  • Experience with infrastructure-as-code tools (Ansible/Terraform)
  • Experience with GitOps (Helm, ArgoCD)
  • Strong grasp of RDMA, InfiniBand, and GPUDirect technologies
  • Familiarity with deep learning frameworks such as PyTorch and TensorFlow
  • Familiarity in at least one cloud platform: AWS, Azure or GCP

If youre a natural problem-solver with a passion for continuous learning, wed love to hear from you.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

#J-18808-Ljbffr

Create a job alert for this search

Site Reliability Engineer, AI/ML Infrastructure • Toronto, ON, CA

Similar jobs

Sr. Site Reliability Administrator

OpenTextRichmond Hill, York Region, CA
Full-time

Opentext - The Information Company.OpenText is a global leader in information management, where innovation, creativity, and collaboration are the key components of our corporate culture.As a member... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, ON, CA
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office].The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Senior Site Reliability Engineer – Cloud & Automation Lead

Tecsys Inc.Toronto, ON, CA
Full-time

A leading supply chain solutions provider is seeking a Site Reliability Engineer to optimize and ensure the reliability of their cloud infrastructure across AWS and Kubernetes.This role emphasizes ... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsToronto, ON, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Site Reliability Engineer for Cloud Infrastructure Management

NewtonToronto, ON, CA
Full-time

Be a pivotal Site Reliability Engineer focused on improving infrastructure resilience and reliability.Collaborate remotely to drive operational success and enhance system performance in a dynamic e... Show more

 • Promoted

AI/ML Solutions Architect for Energy Systems

GE VernovaMarkham, York Region, CA
Full-time

Shape the future of energy with your expertise as an AI/ML Solutions Architect.Drive the design and deployment of machine learning and generative AI applications tailored for grid automation.This h... Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto, ON, CA
Permanent

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

IBM Site Reliability Engineering Expert

LeadingtalentMarkham, ON, CA
Full-time

Step into a career as a Site Reliability Engineer at IBM, focused on enhancing system reliability and performance.Engage directly with production systems and optimize customer experience.In this ro... Show more

 • Promoted

ML Engineer (Hybrid) — Pipelines & MLOps

Aviva CanadaMarkham, York Region, CA
Full-time

A leading insurance company in Markham is seeking an Intermediate Machine Learning Engineer to design and implement ML pipelines and collaborate with cross-functional teams.The ideal candidate will... Show more

 • Promoted

Site Reliability Engineer, Inference Infrastructure

CohereToronto, Ontario, Canada
Full-time

Cohere is the leading security-first enterprise AI company.We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training ... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificToronto, ON, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more