Talent.com
eTeam
Cloud DevOps EngineereTeam • Toronto, ON
Cloud DevOps Engineer

Cloud DevOps Engineer

eTeam • Toronto, ON
14 hours ago
Job type
  • Full-time
Job description
Role Name: Cloud DevOps Engineer
Work site: Toronto (Onsite)
Senior DevOps / Site Reliability Engineer – AI Platform

Role Overview:

We are looking for a Senior DevOps / Site Reliability Engineerto help build, operate, monitor, and scale our new enterprise AI platform.
This role will be responsible for the DevOps and reliability capabilities supporting the platform across all environments, including Development, QA, and Production. The environment is currently focused and manageable, consisting of approximately – containers, but is expected to grow as the AI platform expands in and beyond.
This is an opportunity to work with a collaborative team on a new platform where the focus is not simply deploying infrastructure. The most important part of the role will be creating intelligent dashboards, monitoring, alerting, automated scaling, and AI-assisted platform management.
The platform is hosted primarily in Microsoft Azure, with limited exposure to AWS.

Key Responsibilities:
Azure Infrastructure and Platform Deployment:
  • Design, deploy, configure, and maintain infrastructure within Microsoft Azure.
  • Deploy and manage virtual machines, containers, Kubernetes clusters, networking, storage, and supporting platform services.
  • Support infrastructure across Development, QA, and Production environments.
  • Establish repeatable and reliable deployment processes using Infrastructure as Code and CI/CD automation.
  • Maintain secure, resilient, and appropriately sized platform environments.

Kubernetes and Container Management:
  • Deploy, configure, and operate containerized applications using Kubernetes.
  • Manage container lifecycle, configuration, secrets, networking, storage, and application dependencies.
  • Monitor container and cluster health, resource consumption, capacity, and performance.
  • Troubleshoot deployment, networking, configuration, and runtime issues.
  • Establish appropriate standards for container deployment and Kubernetes operations.

Performance, Load Management, and Scaling:
  • Monitor platform demand, workload patterns, resource utilization, and application performance.
  • Configure horizontal and vertical scaling policies for containers and supporting infrastructure.
  • Develop intelligent scaling approaches based on workload, queue depth, response time, resource utilization, and business demand.
  • Conduct capacity planning and identify potential performance bottlenecks before they affect production.
  • Help introduce predictive or AI-assisted scaling and platform management capabilities.

Dashboards and Platform Visibility:
  • Design and build advanced operational dashboards using tools such as Grafana, Kibana, Azure Monitor, Application Insights, and similar technologies.
  • Create clear executive, operational, application, and infrastructure views of platform health.
  • Build dashboards covering availability, performance, capacity, errors, latency, traffic, container health, AI workloads, and service dependencies.
  • Establish meaningful service-level indicators, service-level objectives, and reliability metrics.
  • Continuously improve dashboards so that issues, trends, and risks can be quickly identified.
  • Advanced dashboard design and dashboard-building experience is a core requirement for this role.

Monitoring and Alerting:
  • Implement monitoring and alerting across infrastructure, applications, containers, integrations, and AI platform services.
  • Configure actionable alerts that identify real production risks while minimizing unnecessary alert noise.
  • Establish thresholds, anomaly detection, health checks, synthetic monitoring, and automated remediation where appropriate.
  • Create operational runbooks and troubleshooting guidance.
  • Work with development and architecture teams to improve platform observability.

Production Reliability and Support:
  • Support the stability, availability, and operational readiness of the production AI platform.
  • Investigate and resolve platform, deployment, infrastructure, monitoring, and performance issues.
  • Participate in root-cause analysis and implement preventative improvements.
  • Ensure that production support processes, documentation, and escalation paths are established before platform usage increases.
  • Provide very light production support during , with no regular after-hours support currently anticipated.
  • Help prepare the operating model for increased platform adoption and support requirements expected in .

Required Qualifications:
  • Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure, or platform engineering.
  • Advanced hands-on experience with Microsoft Azure.
  • Strong experience deploying and operating Kubernetes environments.
  • Strong knowledge of containerization technologies such as Docker.
  • Experience deploying and supporting containerized applications in Development, QA, and Production environments.
  • Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights, or comparable tools.
  • Strong experience implementing monitoring, observability, logging, alerting, and operational health checks.
  • Experience managing application load, infrastructure capacity, performance, and automated scaling.
  • Experience with CI/CD pipelines and automated application deployment.
  • Experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates.
  • Strong troubleshooting skills across applications, containers, infrastructure, networking, and cloud services.
  • Ability to work independently while collaborating closely with developers, architects, AI engineers, and platform stakeholders.

Preferred Qualifications:
  • Experience supporting AI, machine learning, data, or high-compute platforms.
  • Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues, or model performance.
  • Experience implementing automated remediation, predictive monitoring, or AI-assisted platform operations.
  • Familiarity with AWS services and cloud operations.
  • Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus, or similar observability technologies.
  • Experience defining service-level indicators, service-level objectives, and reliability standards.
  • Experience with security, identity, secrets management, and cloud governance within Azure.

What Makes This Role Different:
  • This is not a large-scale, high-pressure production support environment. The initial platform footprint is relatively focused, with approximately – containers and very limited production support expected during .
  • The role offers the opportunity to establish the platform correctly from the beginning, introduce modern DevOps and SRE practices, and experiment with intelligent monitoring, automated scaling, advanced dashboards, and AI-assisted platform management.
  • As platform adoption increases, the responsibilities and operational scope are expected to grow throughout .

Ideal Candidate:
  • The ideal candidate is a senior, hands-on DevOps or SRE professional who enjoys building reliable cloud platforms but is equally interested in observability, intelligent automation, dashboards, and operational innovation.
  • They should be comfortable working in a new and evolving environment, establishing practical standards without unnecessary complexity, and helping a collaborative team build a modern enterprise AI platform.
Create a job alert for this search

Cloud DevOps Engineer • Toronto, ON

Similar jobs

DevOps Cloud Engineer Needed ASAP

Source CodeToronto, ON, CA
Temporary

Join a cutting-edge initiative as a DevOps Cloud Engineer for a 9-month contract.Utilize your expertise in cloud computing and DevOps principles to develop groundbreaking digital solutions.Candidat... Show more

 • Promoted

DevOps Cloud Engineer – GCP & Kubernetes

DelpathToronto
Full-time

DevOps Cloud Engineer – GCP and Kubernetes.Cloud Engineering – The client has embarked on the journey to modernize both development practices and tools.One of the main areas of transformation is th... Show more

 • Promoted

Cloud Platform Engineer — AWS, Kubernetes & DevOps

CI FinancialToronto
Full-time

A leading financial services firm in Toronto is looking for a Cloud Platform Engineer to work within a cross-functional team.This role involves designing and implementing robust cloud solutions usi... Show more

 • Promoted

Senior DevOps Engineer

EQ BankToronto, ON, CA
Full-time

As we continue to scale our team, candidates selected for our comprehensive interview process may be considered for a different level — either higher or lower — based on their interview performance... Show more

 • Promoted

Hybrid Cloud DevOps Engineer | Kubernetes & IAM

Rubicon PathToronto, ON, CA
Full-time

A leading technology consultancy in Toronto is seeking a DevOps/Cloud Engineer to manage and optimize cloud environments while ensuring compliance with public sector regulations.This role involves ... Show more

 • Promoted

Senior Cloud & DevOps Engineer - Remote | Unlimited PTO

Lazer TechnologiesToronto, ON, CA
Remote
Full-time

A world-class digital product studio is seeking a Senior Infrastructure/DevOps Engineer to support a remote-first team.The ideal candidate will have over 5 years of experience, mastery in Docker an... Show more

 • Promoted

Senior Cloud & DevOps Engineer - IaC, Kubernetes, CI/CD

LazerToronto, ON, CA
Full-time

Apple, Google, Coinbase, and more.With our product experience, we have designed, engineered, and grown products.Clients seek out our help because we have the talent to deeply understand their needs... Show more

 • Promoted

Lead Cloud DevSec Ops Engineer

Highbrow LLCToronto, ON, CA
Full-time

Lead Cloud DevSec Ops Engineer.DevOps, Python, PowerShell, Rego, Azure, GCP.Ensure that all cloud solutions follow internally defined security and compliance controls.Implement the enterprise cloud... Show more

 • Promoted

Azure DevOps Engineer

Swoonmarkham, york region, Canada
Full-time

This range is provided by Swoon.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Direct message the job poster from Swoon.Technical Recruiter @ S... Show more

 • Promoted

Remote DevOps Engineer for Cloud, IaC & Automation

Modaxo Inc.Toronto, ON, CA
Remote
Full-time

A leading technology organization is seeking a DevOps Engineer to manage cloud infrastructure and enhance system operations.You will work across multiple business units, ensuring operational excell... Show more

 • Promoted

Cloud DevOps Engineer: Terraform, Kubernetes, CI/CD

PaymentusRichmond Hill, York Region, CA
Full-time

A tech company in York Region is looking for a DevOps Engineer responsible for supporting and monitoring cloud deployments.The successful candidate will collaborate closely with Development and QA ... Show more

 • Promoted

DevOps Engineer

TechDoQuestToronto, ON, CA
Full-time

This job posting is for an existing, active vacancy and we are looking to hire a DevOps Engineer who has expertise/experience working in AEM Cloud Manager.We are looking for a DevOps Engineer suppo... Show more

 • Promoted

Cloud & DevOps Software Engineer — Hybrid (Toronto)

4FMV IncToronto, ON, CA
Full-time

A technology services company in Toronto is seeking a Software Developer to join a hybrid role.This position involves working across software development and cloud infrastructure, focusing on desig... Show more

 • Promoted

Lead AWS Cloud DevOps Engineer (Remote - Namer)

JobgetherToronto, ON, CA
Remote
Full-time

Lead AWS Cloud DevOps Engineer.Location: North America (Remote).As a Lead AWS Cloud DevOps Engineer, you will oversee and build cloud infrastructure that powers mission‑critical platforms, ensuring... Show more

 • Promoted

DevOps Engineer

TekWissen ®Markham, ON, CA
Full-time

TekWissen is a global workforce management provider headquartered in Ann Arbor, Michigan that offers strategic talent solutions to our clients world-wide.This Client is an American multinational se... Show more

 • Promoted

Senior DevOps Engineer Enhancing Cloud Infrastructure and Operations

ScotiabankToronto
Full-time

Shape a transformative data platform as a Senior DevOps Engineer.Utilize your expertise in cloud infrastructure, CI/CD, and operational support for a high-performing team focused on compliance and ... Show more

 • Promoted

DevOps Engineer

VySystemsToronto, ON, CA
Full-time

Senior Site Reliability Engineer (Remote First).Direct message the job poster from VySystems.IT disciplines, including technical architecture, network management, application development, middlewar... Show more

 • Promoted

Senior DevOps Engineer in Cloud Environments

one37Toronto, ON, CA
Full-time

Join us as a Senior DevOps Engineer, integral to enhancing our Web3.Collaborate in a fast-paced, innovative space alongside product management and engineering teams.This role entails providing depl... Show more

 • Promoted

Senior Multicloud DevOps Engineer - Remote

LumenaltaToronto, ON, CA
Remote
Full-time

A leading software solutions company in Vancouver is seeking a skilled cloud engineer to design and implement scalable cloud solutions.The ideal candidate will have over 6 years of experience with ... Show more

 • Promoted

DevOps Engineer

OceanMD, a WELLSTAR CompanyToronto, ON, CA
Full-time

Join us as we change healthcare for the better.OceanMD, a WELLSTAR Company, is the leading provider of EMR-integrated Patient Engagement and eReferral tools in Canada, playing a critical role in mi... Show more