Talent.com

Reliability engineer Jobs in Etobicoke, ON

Create a job alert for this search

Reliability engineer • etobicoke on

Last updated: 1 day ago

Senior Site Reliability Engineer - AEM - Content Delivery Network

astra north infoteckToronto, ON, ca
Full-time

Senior Site Reliability Engineer - AEM - Content Delivery Network.Application Support & Incident Management.Own end-to-end monitoring of the controlled surface: CDN and edge configuration, DNS,... Show more

Reliability Engineer

kinross goldToronto, ON, CA
Full-time

Location: Downtown Toronto (outside Union Station – TTC & GO accessible).Founded in 1993, Kinross is a Canadian-based senior gold mining company with operations and projects in the United State... Show more

Senior RF Engineer - Network Engineer

lancesoftMississauga, ON, CA
CA$50.00–CA$53.00 hourly
Full-time +1
Quick Apply

We are seeking an experienced Senior RF Engineer to support the planning, design, compliance, quality review, and delivery of Radio Access Network solutions.Experience with Canadian wireless-networ... Show more

Data Engineer

zurich insuranceToronto, ON, CA
Full-time

Are you looking for a caring, collaborative, values-driven workplace with inspiring teammates and leaders? Do you have the ambition and desire to be the best and thrive at the most impactful global... Show more

Solution Engineer

prophixEtobicoke, ON, CA
CA$150,000.00 yearly
Full-time
Quick Apply

See what you can do with Prophix.Prophix helps finance teams work with greater flexibility and confidence through Prophix One™, our Financial Performance Platform.We bring planning, reporting, and ... Show more

Solutions Engineer

stan aiToronto, Ontario, Canada
CA$80,000.00–CA$110,000.00 yearly
Full-time
Quick Apply

This is a solutions engineer role, you own the technical health of your deployments end to end: when the AI gets something wrong, you find out why — digging through conversation logs, refining prom... Show more

Civil Engineer

hatchMississauga, ON, CA
Full-time

Join a company that is passionately committed to the pursuit of a better world through positive change.With more than 70 years of business and technical expertise in.With practical solutions that a... Show more

Staff Site Reliability Engineer

case wareToronto, Ontario, Canada, M9W
CA$140,000.00 yearly
Full-time +1

Staff Site Reliability Engineer.Caseware is one of Canada's original Fintech companies, having led the global audit and accounting software industry for over 30 years, with more than 500,000 users ... Show more

Manager, Network Reliability and Resiliency

service nowToronto, Ontario, Canada
CA$125,700.00 yearly
Full-time +1

Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment.The screeni... Show more

Engineer

toronto hydroToronto, ON, CA
Full-time

Target Variable Performance Pa.The salary range shown above reflects the expected compensation for this position.The final salary offered will be determined based on a holistic assessment of the ca... Show more

Site Reliability Engineer

totem recruteur de talentToronto, ON, CA
Permanent

Schedule: 40 hours/week – 100% remote work.We are looking for an experienced.Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infr... Show more

Senior Site Reliability Engineer

i manageToronto, ON, CA
Full-time
Quick Apply

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams -- SRE teams are anchored ... Show more

Senior Service Reliability Engineer

fitch solutionsToronto, ON, CA
CA$120,000.00–CA$150,000.00 yearly
Full-time

Fitch Group is currently seeking a Senior Service Reliability Engineer to embed with Fitch Ratings development squads in Toronto and to partner with developers across Chicago, London, Manchester an... Show more

Project Engineer

alstomToronto, ON, CA
CA$90,000.00–CA$120,000.00 yearly
Full-time

At Alstom, we understand transport networks and what moves people.From high-speed trains, metros, monorails, and trams, to turnkey systems, services, infrastructure, signalling and digital mobility... Show more

Site Reliability Engineer

royal bank canadaToronto, Ontario
Full-time

The SRE will be responsible for assisting in the development, implementation and support of Site Reliability Engineering solutions for all applications across a line of business within CNB (City Na... Show more

Staff Site Reliability Engineer/ Azure

motion recruitmentToronto, ON, Canada
Full-time

Join a growing financial technology organization where your engineering expertise will make a meaningful difference in how schools across North America manage their financial operations.As a Staff ... Show more

Cybersecurity Engineer

xanaduToronto, Ontario, CA
CA$80,000.00–CA$100,000.00 yearly
Full-time

Xanadu’s mission is to build quantum computers that are useful and available to peopleeverywhere.At Xanadu, we are learners, innovators, researchers, collaborators and problem solvers.We are creati... Show more

Site Reliability Engineer- TDJP00058343

randstad canadaToronto, Ontario, CA
Full-time +2
Quick Apply

Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team.In this engineering-focus... Show more

Reliability Expert - Fully Remote | Upto $120/hr

mercorToronto, Ontario, Canada
CA$80.00 hourly
Remote
Part-time
Quick Apply

Headquartered in San Francisco, our investors include.Incident management / reliability / SRE Evaluator.Evaluate AI-generated artifacts against domain-specific quality rubrics.Identify factual, aes... Show more

Structural Engineer

actalentMississauga, Ontario, Canada
Full-time +1

Job Title: Structural Engineer.Join our Civil Engineering Department engaged in nuclear civil engineering work involving structural analysis, civil design, and ageing management.You will support th... Show more

People also ask
Senior Site Reliability Engineer - AEM - Content Delivery Network

Senior Site Reliability Engineer - AEM - Content Delivery Network

astra north infoteckToronto, ON, ca
1 day ago
Job type
  • Full-time
Job description

Job Description

Senior Site Reliability Engineer - AEM - Content Delivery Network


Toronto- 4 Days WFO

ABOUT THE ROLE

1. Application Support & Incident Management

• Own end-to-end monitoring of the controlled surface: CDN and edge configuration, DNS, certificates, cache and invalidation health, and every third-party integration on the page – Search, Consent Management, Analytics, Personalization and AI services and many more to come.

• Build and run synthetic monitoring from outside the bank network, per template, per language, because internal-only monitoring cannot see the CDN, DNS and certificate failures this architecture is most exposed to.

• Run smoke testing of dependent interfaces on every change and maintain the automation packs that do it.

• Participate in the shared on-call rotation as the platform’s subject-matter escalation, and lead incident management for customer-facing events.

• Own the vendor's escalation path: severity mapping between vendor and internal incident scales, named contacts, evidence capture, and holding the vendor to its commitment during an event.

• Handle a class of incident that does not exist on traditional platforms – content published but not visible, invalidation failure, and authoring-source outages – and make those diagnosable by the service desk rather than by you.

2. Change and Release Reliability

• Design and operate change management for the platform where the Git repository is production: reconcile a merge-to-main deployment model with change control, so that every production change carries an approved record with stalling delivery.

• Own the release pipeline as a production control – branch protection, required checks, lint, performance, and secret-screening gates – and the evidence that they are enforced.

• Own rollback: revert, republish and purge, rehearsed end to end with a measured recovery time and a named authority who can call it without convening a meeting.

• Treat content publishing as a routine process: hundreds of production changes made by content authors, needing approval evidence, attribution and retention trail.

• Represent the platform at change advisory board, and own the freeze calendar interaction and release notes.

3. Business Continuity and Resilience

• Own the recovery obligation. The vendor operates delivery resiliently, but customers restore their own content from source version history rather than vendor backups – so the content source, the Git repository and the CDN configuration are the recovery surface, each needing a tested restore.

• Hold CDN and edge configuration as code so that a lost or corrupted property is a redeploy rather than an outage with no runbook.

• Define RTO and RPO with the business against the application criticality tier, document the DR exercise plan, and execute the testing – including failure modes you can actually cause: certificate expiry, invalidation failure, WAF misconfiguration, content source unavailability, and repository compromise.

• Maintain the operational resilience evidence a regulator expects for a material third-party technology arrangement, and keep the platform exit and portability plan current.

4. Reliability & Performance Engineering

• Set and defend service level objectives for both availability and page performance. Define Core Web Vitals thresholds per template, run them on an error budget, and report against them.

• Build the observability practice from the telemetry that exists; real user monitoring on the production domains, CDN access logs streamed to enterprise SIEM as the log source of record, and external synthetics. There is no origin server log – designing around that constraint is part of the job.

• Own third-party scripts and tag governance as a reliability control. Tags are the dominant cause of performance regressions and are added by teams outside engineering change control; you will define the approval route, measure each tag’s cost and enforce the budget.

• Own capacity and cost where they still exist: CDN egress, asset storage and processing, media delivery and any hosted APIs behind the page. Capacity planning here is a financial operations discipline, not a server-sized one.

• Publish the reliability and performance reporting that the business, risk and technology leadership use.

5. Compliance and Control Evidence

• Evidence controls on a platform the organization does not operate – which is harder than evidencing your own, and is where a meaningful share of the role’s effort sits.

• Own log ingestion into SIEM with the agreed retention, access recertification across the repository, admin console, content source and CDN, and the audit evidence pack.

• Support privacy, operational risks, control assessment and third-party risk processes with operational evidence and maintain alignment to regulatory expectations for technology, cyber and third-party risks.

• Keep the configuration management database, support model and assignment groups accurate as the platform estate grows.

WHAT WILL YOU DO?

This is a build-then-run role. Roughly half of the first year is establishing a reliability practice that does not exist yet.

• The observability stack: Real User Monitoring, Core Web Vitals dashboards and alerting, CDN log ingestion, and external synthetics.

• CDN and Edge configuration as Code, with a tested restore.

• The operations runbook, incident playbook, operational level agreement and vendor escalation matrix.

• The change model that reconciles Git-based deployments and continuous content publishing with change controls.

• The first disaster recovery exercise and the first rehearsed, measured rollback.

• Service level objectives agreed with the business, and the reporting that holds the platform to them.

WHAT DO YOU NEED TO SUCCEED?

Must have

• Substantial hands-on experience operating a high-traffic public website behind an enterprise content delivery network. Depth in CDN configuration – origin and cache behaviour, invalidation, edge logic, TLS and DNS – is the single most important qualification. Akamai and Cloudflare experience is an advantage.

• Practical web application firewall experience, including tuning false positives against production-like traffic before enforcement, and bot management that protects the site without blocking the crawlers you need.

• A real observability practice: defining service level objectives and error budgets, and building monitoring from log, real-user and synthetic sources rather than from an agent on a server.

• Web performance engineering – Core Web Vitals, load and rendering behaviour, and the ability to read a waterfall and attribute a regression to a specific script.

• Comfortable with front-end technology: This platform ships JavaScript and CSS to the browser with no server tier; you cannot reason about its reliability without reading and understanding it.

• Git-based release engineering and CI/CD as a production control, including infrastructure and configuration as code.

• Incident command on customer-facing services, and the discipline to produce evidence during an event, not after it.

• Working effectively in a regulated environment – change control, audit evidence, access management and third-party risk – without treating it as an obstacle.

Nice-to-have

• Experience operating a vendor-run or SaaS-delivered platform, where reliability means instrumenting, escalating and holding a supplier accountable rather than fixing the tier yourself.

• Adobe Experience Manager exposure, particularly Edge Delivery Services and Assets as a Cloud Service.

• Financial services or another regulated sector.

• Bilingual delivery – operating a site that must meet the same standard in English and French.

• Accessibility and Search Engine Optimization literacy sufficient to recognize when a reliability decision creates a compliance or discoverability problem.

• Automation in Python, Java or JavaScript, and a preference for encoding a runbook rather than writing one.






Requirements
Java