Talent.com
Chelsea Avondale
Remote Reliability EngineerChelsea Avondale • Toronto, Ontario
Remote Reliability Engineer

Remote Reliability Engineer

Chelsea Avondale • Toronto, Ontario
30+ days ago
Job type
  • Full-time
  • Remote
Job description

Chelsea Avondale is the world’s most cutting-edge home insurance group. We have developed sophisticated risk modeling and insurance pricing technologies for home insurance and deploy that technology through our own insurance company.


Our team consists of some of the brightest minds in insurance, software development, finance, and operations. Our group includes our scientific research & engineering division (Skynet Software) and Canadian property & casualty insurance company (Max Insurance).


Together, our group is transforming the Canadian and global insurance landscape.


JOB DESCRIPTION:


Chelsea Avondale is looking for a Reliability Engineer with a background in infrastructure system engineering to support the growth of a secure, dynamic, and scalable IT environment across the group. Our business is going through rapid growth, and it is essential that our systems infrastructure keeps pace.


The Reliability Engineer will play a crucial role in ensuring the reliability, scalability, and performance of our systems, enabling the continuous delivery of our products and services. They will be accountable for ensuring overall availability, as well as enhancing Engineering teams’ capability to design, build and operate robust systems at scale.


This position is ideal for candidates who have an extraordinary sense of responsibility and are not afraid to roll up their sleeves. Our IT environment is not toolkit rich. What we are NOT looking for is someone who wants to take months installing a large number of tools from their preferred toolkit. We take pride in maintaining a fundamental stack of technologies, much of it in Python, and we are looking for someone who shares this mentality.


If you are someone who thrives in a high-performance culture and is eager for work that is both challenging and constantly evolving, this role is perfect for you. We strongly encourage and help our team members to improve and enhance their personal skill sets within our organization. On your journey with us, you will have the ability to learn and grow rapidly, taking on more responsibilities.


RESPONSIBILITIES:



  • Play an integral role in the design, implementation & maintenance of AWS cloud server environments.

  • Design, implement, and maintain robust monitoring and alerting systems in Python to detect and respond to incidents in a timely manner.

  • Collaborate with cross-functional teams to enhance reliability of our systems and services.

  • Design, configure, deploy, and maintain infrastructure on AWS using best practices and industry standards.

  • Conduct post-incident analysis to identify root causes, implement corrective actions, and prevent similar issues in the future.

  • Assist in capacity planning & optimize services to provide scalable, stable, & secure systems.

  • Implement high availability and disaster recovery solutions to provide data redundancy, resilience, and data loss prevention.

  • Assist with the implementation of select network engineering solutions including firewalls, load balancing, VPNs & LANs, where necessary.


PREFERRED EXPERIENCE & SKILLS:



  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or related field.

  • 1+ years of experience as a Reliability Engineer or similar role, with a focus on maintaining high-performance, scalable, and reliable web systems.

  • We also encourage highly motivated new grads to apply.

  • Hands-on experience with AWS cloud environments – instances, CloudWatch, EFS, etc.

  • Proficiency at Python is a must.

  • Experience using NGINX for reverse proxy, load balancing, and caching.

  • Experience with Unix / Windows server configuration, administration, performance tuning and troubleshooting.

  • Working knowledge of web technologies (web servers, DNS, SSL, Browsers).

  • Working knowledge of web development processes (source control, deployment, etc.).

  • Experience load testing, pen testing, and providing security for cloud resources is beneficial.


Skynet Software welcomes and encourages applications from people with disabilities. Accommodations are available on request for candidates taking part in all aspects of the selection process.

Create a job alert for this search

Remote Reliability Engineer • Toronto, Ontario

Similar jobs

Site Reliability Engineer Role at Magnet Forensics

Magnet ForensicsToronto
Full-time

Take your expertise to the next level as a Senior Site Reliability Engineer with Magnet Forensics.This role involves hands-on AWS and Kubernetes management to uphold our SaaS platform’s reliability... Show more

 • Promoted

Lead Power Platform Reliability Engineer - $127,330 - $236,470 A Year - Remote

ManulifeEast York, Canada
Remote
Full-time

Lead Power Platform Reliability Engineer to enhance service capabilities, design enterprise-level applications, and integrate with key systems.This fully remote role involves stakeholder collaborat... Show more

 • Promoted

Lead Power Platform Reliability Engineer - C$113,260 - C$210,340 A Year - Remote

Manulife John HancockToronto County, Canada
Remote
Full-time

Lead Power Platform Reliability Engineer to enhance service capabilities, develop enterprise-level applications, and ensure seamless integration.Role involves stakeholder collaboration, solution co... Show more

 • Promoted

Senior Site Reliability Engineer I - C$183,000 - C$203,000 A Year - Remote

InstacartToronto County, Canada
Remote
Full-time

Senior Site Reliability Engineer to maintain platform operations, optimize performance, and develop scalable infrastructure.Will lead incident management, monitor systems, and deploy automation tools. Show more

 • Promoted

Site Reliability Engineer

PheedLoopToronto, ON, CA
Full-time

Build the tech behind live events.PheedLoop's mission is to help organizers turn ordinary events into unforgettable experiences with event technology that is bold, intuitive, and built to bring peo... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsToronto, Ontario, Canada
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, Ontario, Canada
Full-time

Working Location: Toronto, ON [Hybrid 2 days a week in office] Role Mandate.The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the rel... Show more

 • Promoted

Senior Site Reliability Engineer

RootlyToronto, Ontario, Canada
Full-time

About Rootly At Rootly, we are on a mission to be the go‑to way companies respond when things go wrong, helping every organization be more reliable.We do this by building an industry‑leading incide... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Software Engineer Site Reliability Engineer - $100,000 - $150,000 A Year - Remote

ThumbtackEast York, Canada
Remote
Full-time

Seeking a Site Reliability Engineer to design, build, and maintain scalable and reliable software systems.Responsibilities include architectural direction, tooling, debugging, capacity planning, an... Show more

 • Promoted

Senior Site Reliability Engineer - C$140,000 - C$155,000 A Year - Remote

McGraw HillToronto, Canada
Remote
Full-time

Seeking a Senior Site Reliability Engineer to build and support reliable, high-capacity systems for learning platforms.Responsibilities include automating cloud infrastructure, optimizing performan... Show more

 • Promoted

Senior Site Reliability Engineer - $111,100 - $166,700 A Year - Remote

ThinkificToronto County, Canada
Remote
Full-time

Senior Site Reliability Engineer to optimize platform infrastructure, focusing on scaling, security, and performance through cloud-native practices, Kubernetes, and AWS. Show more

 • Promoted

Senior Site Reliability Engineer Ii - Remote, Scale-Focused - C$183,000 - C$203,000 A Year - Remote

Leading Grocery Delivery ServiceToronto, Canada
Remote
Full-time

Seeking a Senior Site Reliability Engineer to ensure platform performance, establish incident management, and oversee scalable infrastructure strategies.Requires programming and incident management... Show more

 • Promoted

Site Reliability Engineer

CapgeminiToronto, Ontario, Canada
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Sr. Site Reliability Engineer I - C$184,088 - C$294,540 A Year - Remote

AxonToronto, Canada
Remote
Full-time

Seeking a Senior Site Reliability Engineer to build and operate cloud-native services, focusing on identity and security.Responsibilities include developing platforms, implementing reliability best... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseToronto, ON, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Site Reliability Engineer - Canada Wide - Remote

NewtonToronto, ON, CA
Remote
Full-time

Say hello to Newton! We're changing how Canadians trade crypto.Our goal? To make financial freedom something everyone can achieve.We give our customers the tools and knowledge they need to navigate... Show more

 • Promoted

Site Reliability Engineer

Momentum Financial Services GroupToronto
Full-time

At Momentum Financial Services Group, we help people move forward by reimagining how money works for those who need it most.With more than 40 years of experience, we’re the team behind Money Mart—C... Show more