Talent.com

Reliability Jobs in Old toronto, ON

Create a job alert for this search

Reliability • old toronto on

Last updated: 12 hours ago

Senior Site Reliability Engineer - AEM - Content Delivery Network

astra north infoteckToronto, ON, ca
Full-time

Senior Site Reliability Engineer - AEM - Content Delivery Network.Application Support & Incident Management.Own end-to-end monitoring of the controlled surface: CDN and edge configuration, DNS,... Show more

 • New!

Hiring Sr Support Engineer in Montreal, QC

artechToronto, ON
Full-time

Job Title: Sr Support Engineer.Location: Montreal, QC - Hybrid (2-4 Days WFO).Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboar... Show more

Reliability Engineer

kinross goldToronto, ON, CA
Full-time

Location: Downtown Toronto (outside Union Station – TTC & GO accessible).Founded in 1993, Kinross is a Canadian-based senior gold mining company with operations and projects in the United State... Show more

Software Development Engineer, Ring Cloud Connectivity Org

amazon development centre canada ulcToronto, Ontario, CAN
Full-time

Ring is seeking a Mobile Software Development Engineer to join the Cloud Connectivity organization in Toronto, where you'll build and deliver customer-facing mobile experiences that power how milli... Show more

Maintenance Repair Analyst

seaboard transport groupNorth York, ON, CA
Full-time

The Maintenance Repair Analyst plays a key role in strengthening fleet reliability, safety and performance for heavy-duty trucks and trailers.By connecting field expertise with fleet leadership, th... Show more

AZ Delivery Driver (Local)

qy search advisoryNorth York, Ontario, Canada
CA$28.00 hourly
Full-time

Approximately $77,000/year earnings minimum.Monday to Friday, 7:00 AM to 5:00 PM.No weekends, no overnight routes .North York-based distributor in the flooring and building materials space.They sup... Show more

Manager, Network Reliability and Resiliency

service nowToronto, Ontario, Canada
CA$125,700.00 yearly
Full-time +1

Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment.The screeni... Show more

Site Reliability Engineer

totem recruteur de talentToronto, ON, CA
Permanent

Schedule: 40 hours/week – 100% remote work.We are looking for an experienced.Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infr... Show more

Senior Site Reliability Engineer

i manageToronto, ON, CA
Full-time
Quick Apply

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams -- SRE teams are anchored ... Show more

Sr. Builder - Mobile (Sr. SDE), Ring

amazon development centre canada ulc k03Toronto, Ontario, CAN
Full-time

Ring is redefining how millions of people interact with their homes every single day.As a Senior Builder on this feature team, you'll own and evolve some of the most foundational user experiences i... Show more

Principal AI/ML Engineer

vanguardToronto, ON, CA
Full-time

As a Principal AI Engineer, you will serve as a senior technical leader responsible for transforming state-of-the-art AI research into scalable, production-ready capabilities that create measurable... Show more

Site Reliability Engineer

royal bank canadaToronto, Ontario
Full-time

The SRE will be responsible for assisting in the development, implementation and support of Site Reliability Engineering solutions for all applications across a line of business within CNB (City Na... Show more

Staff Site Reliability Engineer/ Azure

motion recruitmentToronto, ON, Canada
Full-time

Join a growing financial technology organization where your engineering expertise will make a meaningful difference in how schools across North America manage their financial operations.As a Staff ... Show more

Reliability Technician

iiamgoldToronto, ON Mine Site, CA
CA$89,000.00–CA$133,500.00 yearly
Full-time

Reliability Technician-(15503).Innovative, Accountable Mining.IAMGOLD is a Canadian-based gold mining company with operations and development projects across North America and West Africa.With flag... Show more

Site Reliability Engineer- TDJP00058343

randstad canadaToronto, Ontario, CA
Full-time +2
Quick Apply

Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team.In this engineering-focus... Show more

Site Reliability Engineer (SRE)

scotiabankToronto, ON, CA
Full-time

You want to be challenged with complex problem solving taking the learnings forward as continuous improvements.You thrive on supporting critical systems requiring a high level of trust, resilience ... Show more

Reliability Expert - Fully Remote | Upto $120/hr

mercorToronto, Ontario, Canada
CA$80.00 hourly
Remote
Part-time
Quick Apply

Headquartered in San Francisco, our investors include.Incident management / reliability / SRE Evaluator.Evaluate AI-generated artifacts against domain-specific quality rubrics.Identify factual, aes... Show more

Sr. Machine Learning Software Verification Engineer

talentlabToronto, Ontario, Canada
Full-time

AI Software Test / Validation Engineer.Technology / AI / Semiconductor.Our client is a global technology leader developing next-generation AI and machine learning solutions for on-device applicatio... Show more

Mid-Senior Mining Professionals

hire resolve comToronto, ON, CA
Full-time
Quick Apply

Hire Resolve is assisting mining organizations in hiring experienced mining professionals across Canada.This is a multi-role opportunity spanning several functions within the sector, including mine... Show more

AI Infrastructure Engineer

palona aiToronto, ON, CA
Full-time
Quick Apply

Palona’s AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks.Infrastructure is therefore part of the p... Show more

People also ask
Senior Site Reliability Engineer - AEM - Content Delivery Network

Senior Site Reliability Engineer - AEM - Content Delivery Network

astra north infoteckToronto, ON, ca
12 hours ago
Job type
  • Full-time
Job description

Job Description

Senior Site Reliability Engineer - AEM - Content Delivery Network


Toronto- 4 Days WFO

ABOUT THE ROLE

1. Application Support & Incident Management

• Own end-to-end monitoring of the controlled surface: CDN and edge configuration, DNS, certificates, cache and invalidation health, and every third-party integration on the page – Search, Consent Management, Analytics, Personalization and AI services and many more to come.

• Build and run synthetic monitoring from outside the bank network, per template, per language, because internal-only monitoring cannot see the CDN, DNS and certificate failures this architecture is most exposed to.

• Run smoke testing of dependent interfaces on every change and maintain the automation packs that do it.

• Participate in the shared on-call rotation as the platform’s subject-matter escalation, and lead incident management for customer-facing events.

• Own the vendor's escalation path: severity mapping between vendor and internal incident scales, named contacts, evidence capture, and holding the vendor to its commitment during an event.

• Handle a class of incident that does not exist on traditional platforms – content published but not visible, invalidation failure, and authoring-source outages – and make those diagnosable by the service desk rather than by you.

2. Change and Release Reliability

• Design and operate change management for the platform where the Git repository is production: reconcile a merge-to-main deployment model with change control, so that every production change carries an approved record with stalling delivery.

• Own the release pipeline as a production control – branch protection, required checks, lint, performance, and secret-screening gates – and the evidence that they are enforced.

• Own rollback: revert, republish and purge, rehearsed end to end with a measured recovery time and a named authority who can call it without convening a meeting.

• Treat content publishing as a routine process: hundreds of production changes made by content authors, needing approval evidence, attribution and retention trail.

• Represent the platform at change advisory board, and own the freeze calendar interaction and release notes.

3. Business Continuity and Resilience

• Own the recovery obligation. The vendor operates delivery resiliently, but customers restore their own content from source version history rather than vendor backups – so the content source, the Git repository and the CDN configuration are the recovery surface, each needing a tested restore.

• Hold CDN and edge configuration as code so that a lost or corrupted property is a redeploy rather than an outage with no runbook.

• Define RTO and RPO with the business against the application criticality tier, document the DR exercise plan, and execute the testing – including failure modes you can actually cause: certificate expiry, invalidation failure, WAF misconfiguration, content source unavailability, and repository compromise.

• Maintain the operational resilience evidence a regulator expects for a material third-party technology arrangement, and keep the platform exit and portability plan current.

4. Reliability & Performance Engineering

• Set and defend service level objectives for both availability and page performance. Define Core Web Vitals thresholds per template, run them on an error budget, and report against them.

• Build the observability practice from the telemetry that exists; real user monitoring on the production domains, CDN access logs streamed to enterprise SIEM as the log source of record, and external synthetics. There is no origin server log – designing around that constraint is part of the job.

• Own third-party scripts and tag governance as a reliability control. Tags are the dominant cause of performance regressions and are added by teams outside engineering change control; you will define the approval route, measure each tag’s cost and enforce the budget.

• Own capacity and cost where they still exist: CDN egress, asset storage and processing, media delivery and any hosted APIs behind the page. Capacity planning here is a financial operations discipline, not a server-sized one.

• Publish the reliability and performance reporting that the business, risk and technology leadership use.

5. Compliance and Control Evidence

• Evidence controls on a platform the organization does not operate – which is harder than evidencing your own, and is where a meaningful share of the role’s effort sits.

• Own log ingestion into SIEM with the agreed retention, access recertification across the repository, admin console, content source and CDN, and the audit evidence pack.

• Support privacy, operational risks, control assessment and third-party risk processes with operational evidence and maintain alignment to regulatory expectations for technology, cyber and third-party risks.

• Keep the configuration management database, support model and assignment groups accurate as the platform estate grows.

WHAT WILL YOU DO?

This is a build-then-run role. Roughly half of the first year is establishing a reliability practice that does not exist yet.

• The observability stack: Real User Monitoring, Core Web Vitals dashboards and alerting, CDN log ingestion, and external synthetics.

• CDN and Edge configuration as Code, with a tested restore.

• The operations runbook, incident playbook, operational level agreement and vendor escalation matrix.

• The change model that reconciles Git-based deployments and continuous content publishing with change controls.

• The first disaster recovery exercise and the first rehearsed, measured rollback.

• Service level objectives agreed with the business, and the reporting that holds the platform to them.

WHAT DO YOU NEED TO SUCCEED?

Must have

• Substantial hands-on experience operating a high-traffic public website behind an enterprise content delivery network. Depth in CDN configuration – origin and cache behaviour, invalidation, edge logic, TLS and DNS – is the single most important qualification. Akamai and Cloudflare experience is an advantage.

• Practical web application firewall experience, including tuning false positives against production-like traffic before enforcement, and bot management that protects the site without blocking the crawlers you need.

• A real observability practice: defining service level objectives and error budgets, and building monitoring from log, real-user and synthetic sources rather than from an agent on a server.

• Web performance engineering – Core Web Vitals, load and rendering behaviour, and the ability to read a waterfall and attribute a regression to a specific script.

• Comfortable with front-end technology: This platform ships JavaScript and CSS to the browser with no server tier; you cannot reason about its reliability without reading and understanding it.

• Git-based release engineering and CI/CD as a production control, including infrastructure and configuration as code.

• Incident command on customer-facing services, and the discipline to produce evidence during an event, not after it.

• Working effectively in a regulated environment – change control, audit evidence, access management and third-party risk – without treating it as an obstacle.

Nice-to-have

• Experience operating a vendor-run or SaaS-delivered platform, where reliability means instrumenting, escalating and holding a supplier accountable rather than fixing the tier yourself.

• Adobe Experience Manager exposure, particularly Edge Delivery Services and Assets as a Cloud Service.

• Financial services or another regulated sector.

• Bilingual delivery – operating a site that must meet the same standard in English and French.

• Accessibility and Search Engine Optimization literacy sufficient to recognize when a reliability decision creates a compliance or discoverability problem.

• Automation in Python, Java or JavaScript, and a preference for encoding a runbook rather than writing one.






Requirements
Java