Talent.com
Acestack
Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)Acestack • Toronto, ON, Canada
Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Acestack • Toronto, ON, Canada
Il y a 1 jour
Type de contrat
  • Temps plein
  • Quick Apply
Description de poste

Job Title: Site Reliability Engineer (Production Reliability, Azure Operations & Databricks)

Job Type: Full-Time

Location: Toronto, ON (Hybrid)

Job Overview

As an Intermediate Site Reliability Engineer, you will maintain, optimize, and ensure the production reliability of enterprise-level Azure and Databricks platforms. You will focus on platform availability, continuous monitoring, incident response, operational readiness, and platform support while collaborating closely with engineering, security, network, and data teams.

Key Responsibilities

  • Platform Reliability & Support: Monitor and support production Azure and Databricks environments to ensure maximum availability, performance, and operational readiness.

  • Incident Response & On-Call: Respond to production incidents, participate in on-call support rotations, join incident bridges, and execute emergency changes when required.

  • Databricks & Data Services: Support Databricks workspaces, clusters, workflows, user access, and Unity Catalog governance (catalogs, schemas, storage credentials). Maintain integrations with ADLS Gen2, Azure Data Factory, Azure SQL, and Key Vault.

  • Infrastructure & Networking: Troubleshoot Azure infrastructure, including storage accounts, Blob Storage, VNets, NSGs, private endpoints, DNS, and hub-and-spoke connectivity.

  • Observability & Monitoring: Manage alerts, dashboards, and platform health using Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.

  • Root Cause & Maintenance: Perform root cause analysis (RCA), problem management, system patching, upgrades, maintenance, and disaster recovery exercises.

  • Operations & Documentation: Maintain operational runbooks and knowledge articles while tracking tickets and tasks in JIRA and ServiceNow.

Required Technical Qualifications (Must-Haves)

  • Experience: 3+ years supporting Azure production cloud infrastructure and 1+ years supporting Databricks environments.

  • OS Administration: 1+ years of Windows Server administration and 1+ years of Linux administration.

  • Data & Storage: Proven experience supporting Azure Storage services, including ADLS Gen2 and Blob Storage.

  • Networking & Security: Understanding of VNets, NSGs, private endpoints, DNS, routing, Entra ID (Azure AD), RBAC, managed identities, and Azure Key Vault.

  • Monitoring Tools: Hands-on experience with Azure Monitor, Log Analytics, Grafana, Prometheus, Dynatrace, Datadog, or New Relic.

  • ITSM & Operations: Hands-on experience with incident escalation, change management, RCA, JIRA, ServiceNow, and operational runbooks.

Preferred Qualifications (Nice-to-Haves)

  • Operational support experience with Azure SQL and Azure Data Factory (integration runtimes, linked services, orchestration).

  • Understanding of Unity Catalog governance and Disaster Recovery/Business Continuity (RTO/RPO).

  • Exposure to AI/GenAI platforms, Azure OpenAI, MLOps, model endpoints, or RAG services.

  • Experience in cost monitoring, capacity planning, and platform health reporting.

Key Competencies & Soft Skills

  • Strong analytical and calm problem-solving mindset during critical production incidents.

  • Excellent cross-team collaboration skills (working with network, security, and platform teams).

  • Strong documentation skills and customer-focused approach to platform reliability.

Créer une alerte emploi pour cette recherche

Site Reliability Engineer (Production Reliability, Azure Operations & Databricks) • Toronto, ON, Canada

Offres similaires

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto
Temps plein

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Voir plus

 • Offre sponsorisée

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Temps plein

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

DexianToronto, Ontario, Canada
Temps plein

Working Location: Toronto, ON (Hybrid 2 days a week in office).The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

Guidewire SoftwareToronto, Ontario, Canada
Temps plein

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Voir plus

 • Offre sponsorisée

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Télétravail
Temps plein

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Temps plein

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Voir plus

 • Offre sponsorisée

Site Reliability Engineer - Canada Wide - Remote

NewtonToronto, Ontario, Canada
Télétravail
Temps plein

Say hello to Newton! We're changing how Canadians trade crypto.Our goal? To make financial freedom something everyone can achieve.We give our customers the tools and knowledge they need to navigate... Voir plus

 • Offre sponsorisée

Site Reliability Engineer - C$110,000 - C$130,000 A Year

Compass DigitalEast York, Canada
Temps plein

Join Compass Digital as a Site Reliability Engineer to design, build, and automate cloud-native systems using AWS, Go, and TypeScript in a hybrid work environment across Canada. Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

Sage Recruiting Inc.Toronto, Ontario, Canada
Temps plein

This range is provided by Sage Recruiting Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Senior Site Reliability Engineer (Founding Role).A... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Temps plein

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, Ontario, Canada
Temps plein

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Voir plus

 • Offre sponsorisée

Site Reliability Engineer (SRE)

Tangerine BankToronto
Temps plein +1

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Voir plus

 • Offre sponsorisée

IBM Site Reliability Engineering Expert

LeadingtalentMarkham, ON, CA
Temps plein

Step into a career as a Site Reliability Engineer at IBM, focused on enhancing system reliability and performance.Engage directly with production systems and optimize customer experience.In this ro... Voir plus

 • Offre sponsorisée

Site Reliability Engineer

Future Secure AIToronto, Ontario, Canada
Temps plein

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificToronto, ON, CA
Temps plein

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Télétravail
Temps plein

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

iManageToronto, ON, CA
Temps plein

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams – SRE teams are anchored t... Voir plus

 • Offre sponsorisée

Site Reliability Engineer, Inference Infrastructure

CohereToronto, Ontario, Canada
Temps plein

Cohere is the leading security-first enterprise AI company.We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training ... Voir plus

 • Offre sponsorisée

Senior Site Reliability Engineer

MorningstarToronto, Ontario, Canada
Temps plein

About the Team Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment D... Voir plus

 • Offre sponsorisée

Site Reliability Engineer - C$102,700 - C$137,000 A Year

McCain FoodsToronto County, Canada
Temps plein

Seeking a Site Reliability Engineer to ensure software system reliability and availability by designing resilient architectures, automating infrastructure, and optimizing performance in Azure cloud. Voir plus