Talent.com
Astra North Infoteck Inc.
Site Reliability Engineer (SRE) – Dynatrace & AI ObservabilityAstra North Infoteck Inc. • Toronto, Ontario, Canada
Site Reliability Engineer (SRE) – Dynatrace & AI Observability

Site Reliability Engineer (SRE) – Dynatrace & AI Observability

Astra North Infoteck Inc. • Toronto, Ontario, Canada
30+ days ago
Salary
CA$10.00 hourly
Job type
  • Full-time
Job description
Site Reliability Engineer

Location: Toronto ON
Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)
Required Skills

Site Reliability Engineering (SRE)
DevOps
Dynatrace

Role Summary

Design implement and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability performance and reliability across distributed environments.
Leverage Dynatrace Davis AI automation and cloud technologies to enable proactive monitoring intelligent automation and operational excellence.

Role Description

Dynatrace & AI-Driven Observability

Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
Utilize Dynatrace Davis AI for automated root cause analysis anomaly detection event correlation predictive performance insights and alert noise reduction.
Configure and manage OneAgent deployments Smartscape topology mapping service flow and distributed tracing.
Define and monitor SLIs SLOs and user experience metrics.
Build custom dashboards alerts and observability pipelines.
Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
Enable self-healing automation using Dynatrace event triggers and AI-driven insights.

Automation & Configuration Management

Design and implement automation solutions using Ansible.
Automate configuration management application deployments and environment provisioning.
Develop reusable Ansible playbooks and roles for scalable operations.
Automate operational tasks patching compliance processes and remediation workflows.
Integrate Ansible with CI/CD pipelines and monitoring systems.

Cloud & DevOps

Design and manage cloud-native solutions on AWS with exposure to Azure.
Develop infrastructure using Terraform CloudFormation or AWS CDK.
Build and manage CI/CD pipelines using GitHub Actions Jenkins or GitLab CI.
Develop and deploy serverless solutions using AWS Lambda API Gateway and Step Functions.
Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
Deploy and maintain production environments through automated pipelines.
Optimize cloud infrastructure for cost performance and scalability.

Monitoring & Reliability Engineering

Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
Design unified observability across multi-cloud environments.
Implement logging and distributed tracing strategies.
Work with Docker Kubernetes ECS and AKS environments.
Design fault-tolerant highly available and disaster recovery solutions.
Support incident response on-call activities and root cause analysis (RCA).

Required Qualifications

Proven experience with Dynatrace APM Real User Monitoring (RUM) and infrastructure monitoring.
Strong hands-on experience with Dynatrace Davis AI capabilities.
Experience with Ansible for automation and configuration management.
Deep knowledge of AWS services and cloud-native architectures.
Experience with Infrastructure as Code using Terraform CloudFormation or AWS CDK.
Proficiency in Python (boto3) and Bash scripting.
Experience supporting production-scale environments.
Business Analyst experience.
Scrum Master experience.

Nice to Have

Dynatrace Associate or Professional certification.
Experience with Dynatrace APIs and automation.
Experience building self-healing systems using AI-driven triggers.
Familiarity with Prometheus Grafana and the ELK Stack.
Azure cloud experience and certifications.
Experience with GitOps and Platform Engineering.


Required Skills:

Top 3 Required Skills: 1. IBM Financial transaction 2. Payment flow 3. Support Modernization Detailed Job Description: Design develop and maintain applications built on IBM Financial Transaction Manager (FTM) to support core payments processing. Contribute to the development of payment flows supporting transaction processing. Build and support integrations between FTM and upstream/downstream systems using enterprise integration patterns. Participate in the design development testing deployment and production support. Troubleshoot and resolve application and integration issues in a complex regulated environment. Collaborate with architecture QA and operations teams to ensure platform stability scalability and performance. Support modernization initiatives and enhancements to existing payment hub capabilities. Produce clear technical documentation and participate in code reviews and knowledge sharing.


Employment Type : Full Time
Experience: years
Vacancy: 1
Monthly Salary Salary: 10 - 10
Create a job alert for this search

Site Reliability Engineer (SRE) – Dynatrace & AI Observability • Toronto, Ontario, Canada

Similar jobs

Senior Site Reliability Engineer

Guidewire Softwaretoronto, on, Canada
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, ON, CA
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, ON, CA
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office].The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Full-time

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Show more

 • Promoted

Site Reliability Engineer Role at Magnet Forensics

Magnet ForensicsToronto, ON, CA
Full-time

Take your expertise to the next level as a Senior Site Reliability Engineer with Magnet Forensics.This role involves hands-on AWS and Kubernetes management to uphold our SaaS platform’s reliability... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Site Reliability Engineer

KyndrylToronto, ON, CA
Full-time +1

Join to apply for the Site Reliability Engineer role at Kyndryl.Direct message the job poster from Kyndryl.Recruitment & Strategic Staffing @Kyndryl | Partnering with IT Consultants in Financial Se... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Senior Site Reliability Engineer (SRE)

Acquird.ioToronto, ON, CA
Full-time

B2B SaaS company, teams are based out of North America.Role is 95% remote in Toronto (we meetup 1x a month).Must be able to legally work in Canada (visa or sponsorship won't be provided).Our Platfo... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Join Future Secure AI as SRE

Future Secure AIToronto, ON, CA
Full-time

Join Future Secure AI as a Site Reliability Engineer, dedicated to designing and operating robust infrastructure for innovative AI solutions.This role focuses on Kubernetes, cloud technologies, and... Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto, ON, CA
Permanent

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsToronto, ON, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

IBM Site Reliability Engineering Expert

LeadingtalentMarkham, ON, CA
Full-time

Step into a career as a Site Reliability Engineer at IBM, focused on enhancing system reliability and performance.Engage directly with production systems and optimize customer experience.In this ro... Show more

 • Promoted

Site Reliability Engineer, AI/ML Infrastructure

Boson AIToronto, ON, CA
Full-time

We2;re looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters aroundour Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph stor... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

ThinkificToronto, ON, CA
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more