Talent.com
Tubi
Senior Manager, Site Reliability EngineeringTubi • Toronto, ON, Canada
Senior Manager, Site Reliability Engineering

Senior Manager, Site Reliability Engineering

Tubi • Toronto, ON, Canada
30+ days ago
Job type
  • Full-time
Job description

About The Role

Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large‑scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data‑driven decision‑making, blameless learning, and relentless automation.

What You’ll Do

  • Team Leadership & Mentorship:
    • Lead, mentor, and grow a team of Site Reliability Engineers. Foster a culture of innovation and technical excellence where engineers feel empowered to do their best work. Provide personalized coaching, create professional development plans, and guide the careers of senior and emerging talent within the team.
    • Establish equitable, sustainable on‑call practices (including global coverage where applicable) that protect focus time and avoid burnout.
    • Define team rituals – runbook reviews, game days, and incident retros – that reinforce quality and learning.
  • Strategic Planning & Vision: Define and drive the multi‑year technical strategy and vision for Tubi’s observability and automation platforms. Partner with the infra lead to align Tubi’s infrastructure & SRE roadmap. Partner with tech leaders to align the SRE roadmap with business objectives. Champion a data‑driven approach to reliability, using Service Level Objectives (SLOs) and error budgets to facilitate productive conversations about risk and feature velocity.
  • Operational Excellence & Incident Management:
    • Own the end‑to‑end availability, performance, and efficiency of our critical user‑facing services. Evolve our incident response practice to reduce Mean Time to Resolution (MTTR) and Mean Time Between Failures (MTBF). Champion a rigorous, blameless, and data‑driven post‑mortem culture to ensure we learn from both successes and failures, driving engineering teams for systemic fixes and automation to prevent the recurrence of incidents.
    • Streamline and improve our existing processes and practices, and collaborate with other teams to enhance our production release standards by improving current practices.
    • Define and tune a 24×7 on‑call rotation for low noise and fast response; act as executive escalation partner during major incidents.
    • Own disaster‑recovery strategy (playbooks, failover drills, recovery simulations) and track SLO gaps with time‑bound remediations.
  • Financial & Vendor Management: Own the SRE budget, tooling, and headcount. Manage relationships with key third‑party vendors for our observability and SRE related AI platforms, work with infra lead and finance team for contract negotiations and ensure we derive maximum value from our investments.
  • Cross‑Functional Collaboration: Act as a key influencer and strategic partner to leaders in Software Engineering, Product Management, and Infra/Sec. Drive the adoption of SRE best practices and principles throughout the organization, ensuring new services are designed for reliability, scalability, and observability from day one.
  • The AI Mandate: Building the Future of Observability with AI. You will not just manage a team that uses AI; you will lead the charge in building an AI‑native SRE function. This strategic mandate requires a forward‑thinking leader who understands both the potential and the pitfalls of integrating intelligent systems into critical operations. Responsibilities include:
    • AIOps Strategy Development: Developing and executing the strategy for integrating AIOps and machine learning into our observability stack. Your goal will be to move the team from a reactive monitoring posture to one of predictive maintenance and automated anomaly detection, fundamentally changing how we ensure reliability.
    • Accelerating Automation with AI: Championing the effective and responsible use of AI‑assisted coding tools (e.g., Claude Code, Cursor) within the SRE team. You will set the standards and practices to leverage these tools to accelerate the development of automation, operational tooling, and infrastructure code.
    • Building the Business Case: Building the techno‑economic case for new AI tooling, managing vendor relationships, and ensuring the cost‑effective and secure implementation of these powerful systems. You must be able to articulate the ROI of these investments in terms of reduced downtime, improved operational efficiency, and faster incident resolution.
    • Fostering Critical AI Literacy: Fostering a culture that can critically evaluate, debug, and learn from the outputs of AI systems. This involves extending our blameless post‑mortem philosophy to AI‑driven actions and recommendations, ensuring that the team remains in control and understands the "why" behind automated decisions.

Your Background

  • 8+ years of experience in a technical field, with at least a year in an engineering leadership position managing SRE, DevOps, or Production Engineering teams.
  • A deep, principled understanding of SRE tenets, including Service Level Indicators (SLIs), SLOs, error budgets, toil reduction, and capacity planning.
  • Exceptional communication, negotiation, and influencing skills, with the ability to articulate complex technical concepts and strategies to both technical and non‑technical stakeholders at all levels of the organization.
  • A strong technical background as a hands‑on software engineer or site reliability engineer prior to moving into management. Deep knowledge of AWS services (especially networking, IAM, EKS, ALBs/NLBs, Route 53, CloudWatch). Proven experience with Kubernetes in production (EKS preferred), including service exposure, networking, and availability engineering.
  • Hands‑on familiarity with modern SRE tools and technologies, including Infrastructure as Code (e.g., Terraform, Ansible), container orchestration (Kubernetes), observability platforms (e.g., Prometheus, Grafana, Datadog, Splunk), incident tooling (e.g., PagerDuty, FireHydrant), deployment‑safety tooling (e.g., Argo Rollouts, LaunchDarkly), and observability standards (e.g., OpenTelemetry).

Pursuant to local pay disclosure requirements, the pay range for this role, with final offer amount dependent on education, skills, experience, and location, is listed annually below. This role is also eligible for an annual discretionary bonus, long‑term incentive plan, and various benefits including medical/dental/vision, insurance, vacation/paid time off and other benefits in accordance with applicable plan documents.

Toronto, Canada

$164,600 – $235,100 CAD

Benefits

Tubi Media Group is a division of Fox Corporation, and the FOX Employee Benefits summarized here, covers the majority of employee benefits. The following distinctions below outline the differences between the Tubi and FOX benefits:

  • For all salaried employees, in lieu of the FOX Vacation policy, Tubi offers a Flexible Time Off Policy to manage all personal matters.
  • For all full‑time, regular employees, in lieu of FOX Paid Parental Leave, Tubi offers a generous Parental Leave Program, which allows parents twelve (12) weeks of paid bonding leave (top up in Canada) within the first year of birth, adoption, surrogacy, or foster placement of a child in addition to applicable government leave program(s) and FOX’s short‑term disability policy (if applicable). This time is 100% paid through a combination of any applicable government leaves and wage‑replacement programs in addition to contributions made by Tubi.
  • For all full‑time, regular employees, Tubi offers a monthly wellness reimbursement.

About Tubi

Boldly built for every fandom, Tubi is a free streaming service that entertains over 100 million monthly active users. Tubi offers the world’s largest collection of Hollywood movies and TV shows, thousands of creator‑led stories and hundreds of Tubi Originals made for the most passionate fans. Headquartered in San Francisco and founded in 2014, Tubi is part of Tubi Media Group, a division of Fox Corporation.

We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, gender identity, disability, protected veteran status, or any other characteristic protected by law. We will consider for employment qualified applicants with criminal histories consistent with applicable law.

#J-18808-Ljbffr
Create a job alert for this search

Senior Manager, Site Reliability Engineering • Toronto, ON, Canada

Similar jobs

Site Reliability Engineering Senior Manager Role

DawninfotekToronto
Full-time

Shape the future of banking reliability as a Senior Manager in Site Reliability Engineering at a BANK.This role emphasizes incident management and performance optimization strategies.As a Senior Ma... Show more

 • Promoted

Senior Site Reliability Engineer

Guidewire Softwaretoronto, on, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Senior Site Reliability Engineer, Kong Konnect

Kong Inc.toronto, on, Canada
Full-time

Senior Site Reliability Engineer, Kong Konnect.This range is provided by Kong Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Are you ready ... Show more

 • Promoted

Engineering Manager, Transmission Line Projects

Stantec Consulting International Ltd.markham, york region, Canada
Full-time

Manage a growing team of transmission line engineers.Ensure project execution aligns with industry standards in a flexible hybrid environment.As the Engineering Manager for Transmission Line Projec... Show more

 • Promoted

Senior Site Reliability Engineer - Hybrid & Leadership - C$120,000 - C$140,000 A Year

AgriTechToronto County, Canada
Full-time

Senior Site Reliability Engineer needed to lead infrastructure, improve product resilience, and mentor a team in Vancouver.Requires 5+ years of cloud experience (AWS/GCP). Show more

 • Promoted

Site Reliability Engineer

Socket.devtoronto, on, Canada
Full-time

We are seeking a Senior Consultant in Site Reliability Engineering (Network SRE) to lead network-centric reliability practices across the Shared Platform ecosystem.This role focuses on ensuring res... Show more

 • Promoted

Manager, Site Reliability Engineering

Mastercardtoronto, on, Canada
Full-time

Mastercard powers economies and empowers people in 200+ countries and territories worldwide.Together with our customers, we’re helping build a sustainable economy where everyone can prosper.We supp... Show more

 • Promoted

Experienced Project Manager in Demolition

Green Infrastructure Partnersmarkham, on, Canada
Full-time

Take charge of demolition projects with Green Infrastructure Partners Inc.Ensure projects meet high standards while leading teams and managing schedules and budgets.As a Project Manager within Gree... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Senior Manager, Site Reliability Engineering

DawninfotekToronto, Canada
Full-time

Senior Manager, SiteReliabilityEngineering(SRE) Contract to hire for a. Show more

 • Promoted

Senior Site Reliability Engineer

Magnet Forensicstoronto, on, Canada
Full-time

Magnet Forensics is a global leader in the development of digital investigative software that acquires, analyzes, and shares evidence from computers, smartphones, tablets, and IoT-related devices.O... Show more

 • Promoted

Senior Site Reliability Engineer

RootlyToronto, ON, CA
Full-time

At Rootly, we are on a mission to be the go‑to way companies respond when things go wrong, helping every organization be more reliable.We do this by building an industry‑leading incident management... Show more

 • Promoted

Senior Manager, Site Reliability Engineering - C$164,600 - C$235,100 A Year

TubiNorth York, Canada
Full-time

Lead and grow a Site Reliability Engineering team, focusing on platform resilience, automation, and AI integration.Drive technical strategy, operational excellence, and cross-functional collaboration. Show more

 • Promoted

Manager, Site Reliability Engineering - C$140,600 - C$190,600 A Year

Thomson ReutersToronto County, Canada
Full-time

Lead a Site Reliability Engineering team, focusing on system reliability, performance, automation, and DevOps practices.Drive strategic vision, operational excellence, and risk management for cloud... Show more

 • Promoted

Reliability Engineering Manager - $120,000 - $140,000 A Year

Itec GroupEast York, Canada
Full-time

Manages plant reliability, preventative maintenance, and engineering projects to ensure reliable manufacturing operations. Show more

 • Promoted

GIP Site Supervisor - Construction Leadership

03001 GIP - GFLIwhitchurch stouffville, on, Canada
Full-time

Take charge of shoring projects as a Site Supervisor at GIP, ensuring successful execution while prioritizing safety and compliance.Manage daily operations and foster a collaborative work environme... Show more

 • Promoted

Engineering Group Manager - Software Engineering

General Motorsmarkham, on, Canada
Full-time

This posting is for an existing vacancy within the organization and is open to new applications.As part of the application process, Artificial Intelligence will be used in the hiring process for th... Show more

 • Promoted

Senior Engineering Manager, Site Reliability - C$243,000 - C$297,000 A Year

RelayNorth York, Canada
Full-time

Leads the Site Reliability Engineering (SRE) team, defining strategy, improving platform reliability, performance, and resilience, and influencing engineering and product decisions. Show more

 • Promoted

Site Reliability Engineer (Senior or Staff), Deployments

AlleyCorptoronto, on, Canada
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Director, Reliability Engineering

Apotex Inc.toronto, on, Canada
Full-time

Apotex is a Canadian-based global health company.We improve everyday access to affordable, innovative medicines and health products for millions of people worldwide, with a broad portfolio of gener... Show more