Job Description
Senior Site Reliability Engineer - AEM - Content Delivery Network
Toronto- 4 Days WFO
ABOUT THE ROLE
1. Application Support & Incident Management
• Own end-to-end monitoring of the controlled surface: CDN and edge configuration, DNS, certificates, cache and invalidation health, and every third-party integration on the page – Search, Consent Management, Analytics, Personalization and AI services and many more to come.
• Build and run synthetic monitoring from outside the bank network, per template, per language, because internal-only monitoring cannot see the CDN, DNS and certificate failures this architecture is most exposed to.
• Run smoke testing of dependent interfaces on every change and maintain the automation packs that do it.
• Participate in the shared on-call rotation as the platform’s subject-matter escalation, and lead incident management for customer-facing events.
• Own the vendor's escalation path: severity mapping between vendor and internal incident scales, named contacts, evidence capture, and holding the vendor to its commitment during an event.
• Handle a class of incident that does not exist on traditional platforms – content published but not visible, invalidation failure, and authoring-source outages – and make those diagnosable by the service desk rather than by you.
2. Change and Release Reliability
• Design and operate change management for the platform where the Git repository is production: reconcile a merge-to-main deployment model with change control, so that every production change carries an approved record with stalling delivery.
• Own the release pipeline as a production control – branch protection, required checks, lint, performance, and secret-screening gates – and the evidence that they are enforced.
• Own rollback: revert, republish and purge, rehearsed end to end with a measured recovery time and a named authority who can call it without convening a meeting.
• Treat content publishing as a routine process: hundreds of production changes made by content authors, needing approval evidence, attribution and retention trail.
• Represent the platform at change advisory board, and own the freeze calendar interaction and release notes.
3. Business Continuity and Resilience
• Own the recovery obligation. The vendor operates delivery resiliently, but customers restore their own content from source version history rather than vendor backups – so the content source, the Git repository and the CDN configuration are the recovery surface, each needing a tested restore.
• Hold CDN and edge configuration as code so that a lost or corrupted property is a redeploy rather than an outage with no runbook.
• Define RTO and RPO with the business against the application criticality tier, document the DR exercise plan, and execute the testing – including failure modes you can actually cause: certificate expiry, invalidation failure, WAF misconfiguration, content source unavailability, and repository compromise.
• Maintain the operational resilience evidence a regulator expects for a material third-party technology arrangement, and keep the platform exit and portability plan current.
4. Reliability & Performance Engineering
• Set and defend service level objectives for both availability and page performance. Define Core Web Vitals thresholds per template, run them on an error budget, and report against them.
• Build the observability practice from the telemetry that exists; real user monitoring on the production domains, CDN access logs streamed to enterprise SIEM as the log source of record, and external synthetics. There is no origin server log – designing around that constraint is part of the job.
• Own third-party scripts and tag governance as a reliability control. Tags are the dominant cause of performance regressions and are added by teams outside engineering change control; you will define the approval route, measure each tag’s cost and enforce the budget.
• Own capacity and cost where they still exist: CDN egress, asset storage and processing, media delivery and any hosted APIs behind the page. Capacity planning here is a financial operations discipline, not a server-sized one.
• Publish the reliability and performance reporting that the business, risk and technology leadership use.
5. Compliance and Control Evidence
• Evidence controls on a platform the organization does not operate – which is harder than evidencing your own, and is where a meaningful share of the role’s effort sits.
• Own log ingestion into SIEM with the agreed retention, access recertification across the repository, admin console, content source and CDN, and the audit evidence pack.
• Support privacy, operational risks, control assessment and third-party risk processes with operational evidence and maintain alignment to regulatory expectations for technology, cyber and third-party risks.
• Keep the configuration management database, support model and assignment groups accurate as the platform estate grows.
WHAT WILL YOU DO?
This is a build-then-run role. Roughly half of the first year is establishing a reliability practice that does not exist yet.
• The observability stack: Real User Monitoring, Core Web Vitals dashboards and alerting, CDN log ingestion, and external synthetics.
• CDN and Edge configuration as Code, with a tested restore.
• The operations runbook, incident playbook, operational level agreement and vendor escalation matrix.
• The change model that reconciles Git-based deployments and continuous content publishing with change controls.
• The first disaster recovery exercise and the first rehearsed, measured rollback.
• Service level objectives agreed with the business, and the reporting that holds the platform to them.
WHAT DO YOU NEED TO SUCCEED?
Must have
• Substantial hands-on experience operating a high-traffic public website behind an enterprise content delivery network. Depth in CDN configuration – origin and cache behaviour, invalidation, edge logic, TLS and DNS – is the single most important qualification. Akamai and Cloudflare experience is an advantage.
• Practical web application firewall experience, including tuning false positives against production-like traffic before enforcement, and bot management that protects the site without blocking the crawlers you need.
• A real observability practice: defining service level objectives and error budgets, and building monitoring from log, real-user and synthetic sources rather than from an agent on a server.
• Web performance engineering – Core Web Vitals, load and rendering behaviour, and the ability to read a waterfall and attribute a regression to a specific script.
• Comfortable with front-end technology: This platform ships JavaScript and CSS to the browser with no server tier; you cannot reason about its reliability without reading and understanding it.
• Git-based release engineering and CI/CD as a production control, including infrastructure and configuration as code.
• Incident command on customer-facing services, and the discipline to produce evidence during an event, not after it.
• Working effectively in a regulated environment – change control, audit evidence, access management and third-party risk – without treating it as an obstacle.
Nice-to-have
• Experience operating a vendor-run or SaaS-delivered platform, where reliability means instrumenting, escalating and holding a supplier accountable rather than fixing the tier yourself.
• Adobe Experience Manager exposure, particularly Edge Delivery Services and Assets as a Cloud Service.
• Financial services or another regulated sector.
• Bilingual delivery – operating a site that must meet the same standard in English and French.
• Accessibility and Search Engine Optimization literacy sufficient to recognize when a reliability decision creates a compliance or discoverability problem.
• Automation in Python, Java or JavaScript, and a preference for encoding a runbook rather than writing one.
Requirements
Java