Talent.com
PheedLoop
Site Reliability EngineerPheedLoop • Toronto, ON, Canada
Site Reliability Engineer

Site Reliability Engineer

PheedLoop • Toronto, ON, Canada
30+ days ago
Job type
  • Full-time
Job description

Build the tech behind live events

PheedLoop's mission is to help organizers turn ordinary events into unforgettable experiences with event technology that is bold, intuitive, and built to bring people together. From conferences and trade shows to campus events and summits, we help teams run smarter events with tools that feel seamless for both planners and attendees. We bring together check-in kiosks, badge printing, mobile apps, and engagement tools into a connected ecosystem that unifies every part of the event experience. We move fast, we solve real problems, and we care deeply about every organizer trusting us with their biggest moments.

What you’ll build everyday

  • Keep production humming by designing, scaling, and hardening the infrastructure that thousands of users depend on every day. Own uptime, performance, and reliability as first-class product features.
  • Build and evolve CI/CD pipelines, infrastructure-as-code, and automation that take the toil out of deployments and let engineers ship with confidence. Turn manual runbooks into self-healing systems.
  • Partner with backend and frontend teams to translate reliability goals into concrete SLOs, error budgets, and architectural improvements. Drive capacity planning, load testing, and performance tuning across services.
  • Lead incident response when things break, from first page to root cause analysis. Run blameless postmortems, ship the follow-up fixes, and make sure the same fire never starts twice.
  • Instrument the platform end-to-end with logging, metrics, tracing, and alerting so issues get caught before users feel them. Tune signal-to-noise so on-call is sustainable, not soul-crushing.
  • Contribute to peer reviews on infrastructure changes and application code that touches reliability-sensitive paths. Mentor junior engineers and share what you learn with the wider team.
  • Keep a close eye on cloud spend and drive cost optimization initiatives, right-sizing resources, eliminating waste, and architecting for efficiency so reliability and budget both stay in the green.

Skills we are looking for

  • Solid hands‑on experience with a major cloud platform (AWS and GCP) and container orchestration, AWS Fargate strongly preferred. Comfortable operating production workloads at scale.
  • Fluent with infrastructure‑as‑code (Terraform strongly preferred) and configuration management. You treat infrastructure like software: versioned, reviewed, and tested.
  • Strong scripting and automation skills in Python or Bash. Able to jump into application code (Python/Django a plus) to debug issues across the stack, not just the platform layer.
  • Deep familiarity with observability tooling such as CloudWatch. You know the difference between a good alert and a noisy one.
  • Working knowledge of relational databases (Postgres), including replication, backups, query performance, and incident recovery. Exposure to caching layers and message queues is a plus.
  • Comfortable with Git‑based workflows, code review culture, and shipping changes through modern CI/CD systems (GitHub Actions, GitLab CI, CircleCI, or similar).
  • 3+ years in an SRE, DevOps, or production engineering role, with real on‑call experience and a track record of shipping reliability improvements that moved the numbers.
  • Bachelor's degree in Computer Science or related field, or equivalent real‑world experience. Strong written and verbal English communication — you can explain a complex outage to engineers and non‑engineers alike.
  • Calm under pressure, methodical in a crisis, and genuinely curious about how systems fail. Team‑first mindset, sharp problem‑solving instincts, and a bias toward automating the boring stuff away.

Perks

  • Enjoy 100% employer‑paid health coverage, because your well‑being matters.
  • Work from an office directly connected to the TTC subway, making your commute smooth and stress‑free.
  • Join team lunches, learning opportunities, and regular outings that make growth and networking part of the job.
  • Be part of an ambitious, high‑performance culture surrounded by people who love building big things.

Please note: This role is an existing vacancy. We do not use artificial intelligence or automated decision‑making tools in our hiring process.

#J-18808-Ljbffr
Create a job alert for this search

Site Reliability Engineer • Toronto, ON, Canada

Similar jobs

Senior Site Reliability Engineer

Guidewire Softwaretoronto, on, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Site Reliability Engineer Role at Magnet Forensics

Magnet ForensicsToronto, ON, CA
Full-time

Take your expertise to the next level as a Senior Site Reliability Engineer with Magnet Forensics.This role involves hands-on AWS and Kubernetes management to uphold our SaaS platform’s reliability... Show more

 • Promoted

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Full-time

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, Ontario, Canada
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, Ontario, Canada
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office] Role Mandate.The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the rel... Show more

 • Promoted

Site Reliability Engineer - $86,105 - $128,150 A Year

SamsungToronto County, Canada
Full-time

Seeking an SRE/DevOps Engineer to maintain critical systems, focusing on automation, optimization, and operational excellence in cloud-native environments.Collaborates with development teams for se... Show more

 • Promoted

Site Reliability Engineer

KyndrylToronto, ON, CA
Full-time +1

Join to apply for the Site Reliability Engineer role at Kyndryl.Direct message the job poster from Kyndryl.Recruitment & Strategic Staffing @Kyndryl | Partnering with IT Consultants in Financial Se... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Senior Site Reliability Engineer, Kong Konnect

Kong Inc.Toronto, ON, CA
Full-time

Senior Site Reliability Engineer, Kong Konnect.This range is provided by Kong Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Are you ready ... Show more

 • Promoted

Site Reliability Engineer - C$113,400 - C$162,000 A Year

TextNowToronto County, Canada
Full-time

Senior Site Reliability Engineer to design and maintain scalable systems, automate infrastructure, and improve system reliability. Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Site Reliability Engineer

Socket.devToronto
Full-time

We are seeking a Senior Consultant in Site Reliability Engineering (Network SRE) to lead network-centric reliability practices across the Shared Platform ecosystem.This role focuses on ensuring res... Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto
Full-time +1

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Site Reliability Engineer - C$102,700 - C$137,000 A Year

McCain FoodsNorth York, Canada
Full-time

Seeking a Site Reliability Engineer to ensure software system reliability and availability by designing resilient architectures, automating infrastructure, and optimizing performance in Azure cloud. Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsToronto, ON, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer

Momentum Financial Services GroupToronto
Full-time

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more