Job descriptionJob Description
Enterprise Kubernetes SRE (Python, GitOps, API, Container, Cloud,, MongoDB, Postgres)
Toronto, ON - Hybrid (4 Days WFO)
12 months
We are seeking an experienced Site Reliability Engineer to join our Enterprise
Kubernetes Platform team at a leading financial services organization. You'll
be responsible for ensuring the reliability, performance, and scalability of
our enterprise-grade Kubernetes platform that powers mission-critical
applications across the organization.
This role places a strong emphasis on automation, with the expectation that
the successful candidate will continuously identify and eliminate manual toil
through intelligent tooling, self-healing systems, and AI-assisted operational
workflows. You will work alongside platform engineers, DevOps teams, and
application developers to build and maintain a world-class container
orchestration platform.
======================================================================
WHAT YOU'LL DO
======================================================================
PLATFORM RELIABILITY & OPERATIONS
--------------------------------------------------------------------------------
• Ensure 99.9% availability SLA for platform services across 60+ Kubernetes
clusters spanning production, DR, UAT, QA, and development environments
• Manage and operate enterprise Kubernetes distributions across on-premises
and cloud-hosted environments
• Implement and maintain disaster recovery patterns across multi-AZ
architectures and geographically distributed data centres
• Design and execute capacity planning, resource optimization, and cluster
scaling strategies
• Support full cluster lifecycle operations including provisioning, upgrades,
patching, and decommissioning
• Automate cluster health checks and validation workflows for continuous
reliability assurance
• Manage multi-tenant cluster environments with strict isolation and RBAC
enforcement
INCIDENT MANAGEMENT & ON-CALL
--------------------------------------------------------------------------------
• Participate in on-call rotation for platform infrastructure support with
sub-15-minute MTTR targets
• Lead incident response, troubleshooting, and root cause analysis for
platform issues
• Conduct blameless post-incident reviews and implement preventive measures
• Develop and maintain runbooks, troubleshooting guides, and operational
playbooks
• Coordinate with application teams during incidents affecting workloads
• Build automated incident detection and response systems to reduce manual
intervention
• Integrate AI-assisted triage tools for faster incident classification and
resolution