Talent.com
Acestack
Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)Acestack • Toronto, ON, Canada
Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Acestack • Toronto, ON, Canada
21 hours ago
Job type
  • Full-time
  • Quick Apply
Job description

Job Title: Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Job Type: Full-Time

Location: Toronto, ON (Hybrid)

Job Overview

As an Intermediate Site Reliability Engineer, you will maintain, optimize, and ensure the production reliability of enterprise-level Azure and Databricks platforms. You will focus on platform availability, continuous monitoring, incident response, operational readiness, and platform support while collaborating closely with engineering, security, network, and data teams.

Key Responsibilities

  • Platform Reliability & Support: Monitor and support production Azure and Databricks environments to ensure maximum availability, performance, and operational readiness.

  • Incident Response & On-Call: Respond to production incidents, participate in on-call support rotations, join incident bridges, and execute emergency changes when required.

  • Databricks & Data Services: Support Databricks workspaces, clusters, workflows, user access, and Unity Catalog governance (catalogs, schemas, storage credentials). Maintain integrations with ADLS Gen2, Azure Data Factory, Azure SQL, and Key Vault.

  • Infrastructure & Networking: Troubleshoot Azure infrastructure, including storage accounts, Blob Storage, VNets, NSGs, private endpoints, DNS, and hub-and-spoke connectivity.

  • Observability & Monitoring: Manage alerts, dashboards, and platform health using Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.

  • Root Cause & Maintenance: Perform root cause analysis (RCA), problem management, system patching, upgrades, maintenance, and disaster recovery exercises.

  • Operations & Documentation: Maintain operational runbooks and knowledge articles while tracking tickets and tasks in JIRA and ServiceNow.

Required Technical Qualifications (Must-Haves)

  • Experience: 3+ years supporting Azure production cloud infrastructure and 1+ years supporting Databricks environments.

  • OS Administration: 1+ years of Windows Server administration and 1+ years of Linux administration.

  • Data & Storage: Proven experience supporting Azure Storage services, including ADLS Gen2 and Blob Storage.

  • Networking & Security: Understanding of VNets, NSGs, private endpoints, DNS, routing, Entra ID (Azure AD), RBAC, managed identities, and Azure Key Vault.

  • Monitoring Tools: Hands-on experience with Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.

  • ITSM & Operations: Hands-on experience with incident escalation, change management, RCA, JIRA, ServiceNow, and operational runbooks.

Preferred Qualifications (Nice-to-Haves)

  • Operational support experience with Azure SQL and Azure Data Factory (integration runtimes, linked services, orchestration).

  • Understanding of Unity Catalog governance and Disaster Recovery/Business Continuity (RTO/RPO).

  • Exposure to AI/GenAI platforms, Azure OpenAI, MLOps, model endpoints, or RAG services.

  • Experience in cost monitoring, capacity planning, and platform health reporting.

Key Competencies & Soft Skills

  • Strong analytical and calm problem-solving mindset during critical production incidents.

  • Excellent cross-team collaboration skills (working with network, security, and platform teams).

  • Strong documentation skills and customer-focused approach to platform reliability.

Create a job alert for this search

Site Reliability Engineer (Production Reliability, Azure Operations & Databricks) • Toronto, ON, Canada

Similar jobs

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Full-time

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, Ontario, Canada
Full-time

Working Location: Toronto, ON (Hybrid 2 days a week in office).The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Show more

 • Promoted

Senior Site Reliability Engineer

Guidewire SoftwareToronto, Ontario, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Site Reliability Engineer - Canada Wide - Remote

NewtonToronto, Ontario, Canada
Remote
Full-time

Say hello to Newton! We're changing how Canadians trade crypto.Our goal? To make financial freedom something everyone can achieve.We give our customers the tools and knowledge they need to navigate... Show more

 • Promoted

Site Reliability Engineer - C$110,000 - C$130,000 A Year

Compass DigitalEast York, Canada
Full-time

Join Compass Digital as a Site Reliability Engineer to design, build, and automate cloud-native systems using AWS, Go, and TypeScript in a hybrid work environment across Canada. Show more

 • Promoted

Senior Site Reliability Engineer

Sage Recruiting Inc.Toronto, Ontario, Canada
Full-time

This range is provided by Sage Recruiting Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Senior Site Reliability Engineer (Founding Role).A... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, Ontario, Canada
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto
Full-time +1

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

IBM Site Reliability Engineering Expert

LeadingtalentMarkham, ON, CA
Full-time

Step into a career as a Site Reliability Engineer at IBM, focused on enhancing system reliability and performance.Engage directly with production systems and optimize customer experience.In this ro... Show more

 • Promoted

Site Reliability Engineer

Future Secure AIToronto, Ontario, Canada
Full-time

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificToronto, ON, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Site Reliability Engineer

iManageToronto, ON, CA
Full-time

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams – SRE teams are anchored t... Show more

 • Promoted

Site Reliability Engineer, Inference Infrastructure

CohereToronto, Ontario, Canada
Full-time

Cohere is the leading security-first enterprise AI company.We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training ... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto, Ontario, Canada
Full-time

About the Team Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment D... Show more

 • Promoted

Site Reliability Engineer - C$102,700 - C$137,000 A Year

McCain FoodsToronto County, Canada
Full-time

Seeking a Site Reliability Engineer to ensure software system reliability and availability by designing resilient architectures, automating infrastructure, and optimizing performance in Azure cloud. Show more