Talent.com
Mirantis
Senior Site Reliability Engineer (Kubernetes,Mirantis • Toronto, Greater Toronto Area, Canada
Senior Site Reliability Engineer (Kubernetes,

Senior Site Reliability Engineer (Kubernetes,

Mirantis • Toronto, Greater Toronto Area, Canada
9 hours ago
Job type
  • Full-time
  • Remote
Job description
Job Description

We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-thinking cloud development and operations team. In this role, you will contribute to the design, development, and operation of sophisticated cloud-based AI solutions built on the CNCF ecosystem including Kubernetes, running on cutting-edge hardware from leading vendors. This role focuses mainly on deploying AI infrastructure built on NVIDIA-certified hardware, following architecture and implementation designs produced by our engineering team. You will play a pivotal role in ensuring the reliability, security, and performance of container infrastructure, while mentoring team members and Mirantis customers to deliver high-quality software and services. As a senior engineer, you will work closely with stakeholders to define technical strategies, solve complex challenges, and ensure the seamless integration of cloud and software services. This is an excellent opportunity to make a significant impact while driving innovation in a rapidly evolving cloud ecosystem.

Main Responsibilities:

  • Work with geographically distributed international teams on technical challenges and process improvements.

  • Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software.

  • Collaborate with stakeholders to gather and refine technical requirements.

  • Optimize system performance, reliability, and scalability.

  • Troubleshoot, debug, and resolve complex technical issues.

  • Participate in code reviews to maintain high quality standards.

  • Stay up to date with industry trends and best practices in cloud operations and development.

  • Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance.

  • Facilitate knowledge transfer to customers during the delivery phases.


Qualifications

  • 5+ years of professional experience in DevOps, with a strong focus on Cloud, infrastructure technologies and Kubernetes

  • Experience with high-performance data center processing, networking, and storage

  • Exposure to Golang and working knowledge of other programming languages (Python, JavaScript).

  • Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.

  • Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security.

  • Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams.

  • Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight.

  • Excellent written and spoken English.

  • Excellent customer-facing communication skills.

  • A commitment to innovation, continuous learning, and delivering high-quality results.

  • Ability to travel up to 25% if needed, including internationally.

Nice to have

  • Extensive experience in network and/or storage architecture.

  • Experience with high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle or NVIDIA AI Enterprise.

  • Working experience with Openstack

  • Presence in the open source community including upstream contribution and conference presentations.

  • Prior experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware.

Education and Experience:

  • Bachelor's degree in Computer Science or a related field, or equivalent experience.

  • At least 5 years of DevOps or Software Development experience or in a similar role.



Additional Information

What does Mirantis offer you?

- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.


It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to isamoylova@mirantis.com

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

We are a Leader for Container Management in G2 (#2 after AWS)!

Create a job alert for this search

Senior Site Reliability Engineer (Kubernetes, • Toronto, Greater Toronto Area, Canada

Similar jobs

Senior Site Reliability Engineer

Guidewire SoftwareToronto, Ontario, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Full-time

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Show more

 • Promoted

Site Reliability Engineer (Senior or Staff), Deployments

AlleyCorpToronto, Ontario, Canada
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, Ontario, Canada
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office] Role Mandate.The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the rel... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Senior Site Reliability Engineer at Kong

CacheflowToronto, ON, CA
Full-time

Accelerate your career in cloud technologies as a Senior Site Reliability Engineer at Kong.Focus on Managed Gateways, ensuring robust infrastructure and outstanding performance across multiple clou... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsToronto, ON, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Site Reliability Engineer - Canada Wide - Remote

NewtonToronto, Ontario, Canada
Remote
Full-time

Say hello to Newton! We're changing how Canadians trade crypto.Our goal? To make financial freedom something everyone can achieve.We give our customers the tools and knowledge they need to navigate... Show more

 • Promoted

Site Reliability Engineer (Senior or Staff), Deployments

MongoDBToronto, Ontario, Canada
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Senior Site Reliability Engineer

Magnet Forensicstoronto, on, Canada
Full-time

Magnet Forensics is a global leader in the development of digital investigative software that acquires, analyzes, and shares evidence from computers, smartphones, tablets, and IoT-related devices.O... Show more

 • Promoted

PheedLoop Senior Site Reliability Engineer

PheedLoopToronto
Full-time

PheedLoop is seeking a Senior Site Reliability Engineer to maintain and enhance the tech behind live events.This pivotal role prioritizes system reliability and performance.You'll be tasked with de... Show more

 • Promoted

Senior Site Reliability Engineer

iManageToronto, Ontario, Canada
Full-time

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams – SRE teams are anchored t... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, Ontario, Canada
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsToronto, ON, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Site Reliability Engineer

Future Secure AIToronto, Ontario, Canada
Full-time

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificToronto, ON, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Site Reliability Engineer

Robertson & Company Ltd.Toronto, ON, CA
Full-time

Our client is a top financial institution with significant North American holdings.They have operations across most major verticals, including institutional & corporate, wealth management, private ... Show more