Talent.com
Amaris Consulting
Ingénieur(e) Site Reliability (SRE)Amaris Consulting • Toronto, Canada
No longer accepting applications
Ingénieur(e) Site Reliability (SRE)

Ingénieur(e) Site Reliability (SRE)

Amaris Consulting • Toronto, Canada
30+ days ago
Job type
  • Full-time
Job description

Job description

Nous recherchons un(e) Ingénieur(e) Site Reliability (SRE) expérimenté(e) pour soutenir des plateformes de cybersécurité, de gestion des risques liés aux données et de résilience, en garantissant leur fiabilité, disponibilité, performance et visibilité opérationnelle. La personne retenue jouera un rôle clé dans le maintien de la stabilité des environnements de production, l’amélioration de l’observabilité, le soutien à la réponse aux incidents et la fourniture de tableaux de bord et d’analyses opérationnelles destinés aux équipes d’ingénierie, d’exploitation, de gestion des risques et à la direction.

Lieu et mode de travail

  • Lieu : Télétravail partout au Canada

  • Mode de travail : 100 % à distance

Principales responsabilités

  • Maintenir et améliorer la fiabilité, la disponibilité, l’évolutivité et la performance des plateformes de cybersécurité et de l’infrastructure associée.

  • Surveiller l’état de santé des systèmes, identifier les risques opérationnels, répondre aux incidents et assurer une résolution rapide des problèmes ayant un impact sur les services.

  • Instrumenter l’infrastructure, les applications, les API, les pipelines de données, les bases de données et les services cloud afin d’assurer une visibilité opérationnelle de bout en bout.

  • Concevoir, mettre en œuvre et améliorer les solutions de surveillance, d’alerting, de journalisation, de traçage et d’observabilité dans des environnements distribués et cloud natifs.

  • Développer des alertes pertinentes et exploitables afin de réduire le bruit et d’améliorer la qualité du signal pour une détection et une réponse plus rapides.

  • Définir et suivre les indicateurs de fiabilité, notamment les SLI, SLO, SLA, budgets d’erreur, la latence, le débit, les taux d’erreur et les métriques de saturation.

  • Créer et maintenir des tableaux de bord pour les équipes d’ingénierie, d’exploitation, produit, gestion des risques et pour les parties prenantes exécutives.

  • Améliorer continuellement les tableaux de bord exécutifs afin de soutenir les revues de direction sur la santé des services, les tendances d’incidents, les risques opérationnels et les performances de fiabilité.

  • Collaborer avec les équipes d’ingénierie, d’infrastructure, cloud et cybersécurité pour identifier les lacunes de fiabilité et mettre en œuvre des améliorations durables.

  • Participer à la réponse aux incidents, à l’analyse des causes racines, aux revues post-incident et aux activités de gestion des problèmes.

  • Automatiser les tâches opérationnelles, les vérifications de santé, la validation des déploiements, les rapports et les procédures de reprise.

  • Soutenir les processus CI/CD, DevOps et de gestion des mises en production en validant la préparation opérationnelle, la couverture de surveillance, les plans de retour arrière et les exigences de support de production.

  • Contribuer aux initiatives d’ingénierie de résilience, notamment la planification de capacité, l’optimisation des performances, les tests de bascule, la préparation à la reprise après sinistre et les tests de résilience.

  • Veiller à ce que les processus opérationnels et les pratiques d’observabilité soient conformes aux normes de sécurité, de gestion des risques, de conformité et de gouvernance de l’entreprise.

Qualifications requises

  • 10 ans ou plus d’expérience en Site Reliability Engineering, DevOps, ingénierie systèmes, ingénierie d’infrastructure, développement logiciel ou exploitation de production.

  • Solide expérience du support de plateformes hautement disponibles, distribuées, cloud ou critiques pour l’entreprise.

  • Expertise pratique des domaines de la surveillance, des alertes, de la journalisation, des métriques, du traçage, des tableaux de bord et de l’observabilité.

  • Expérience dans l’instrumentation d’applications, de services, d’API, de bases de données, d’infrastructures et de composants cloud.

  • Bonne compréhension des concepts de SLI, SLO, SLA, budgets d’erreur, gestion des incidents, gestion de capacité et préparation opérationnelle.

  • Expérience démontrée dans la conception d’alertes exploitables et dans le soutien à la détection et à la résolution rapides des incidents.

  • Expérience de la création de tableaux de bord opérationnels pour des équipes techniques et des parties prenantes exécutives.

  • Excellentes compétences en programmation ou scripting avec Python, Bash, PowerShell, Java ou des langages similaires.

  • Expérience avec AWS, Azure ou GCP.

  • Expérience avec les outils Infrastructure as Code tels que Terraform ou équivalent.

  • Bonne connaissance des pipelines CI/CD, des pratiques DevOps, des processus de mise en production et des modèles de support de production.

  • Solides compétences en résolution de problèmes dans des environnements distribués, avec des API REST, des plateformes de messagerie et des intégrations de services.

  • Connaissance de PostgreSQL, MSSQL, MongoDB ou d’autres bases de données relationnelles et non relationnelles.

  • Excellentes compétences en communication écrite et orale, avec la capacité d’expliquer des sujets techniques à des publics non techniques et exécutifs.

Atouts appréciés

  • Expérience du support de plateformes de cybersécurité, de gestion des risques, de résilience, de conformité ou de sécurité d’entreprise.

  • Expérience avec des outils d’observabilité tels que Splunk, Grafana, Prometheus, Datadog, Dynatrace, New Relic, Azure Monitor, CloudWatch ou OpenTelemetry.

  • Expérience de la création de tableaux de bord exécutifs sur la santé des services, de scorecards de fiabilité, de rapports de risques opérationnels et de tendances d’incidents.

  • Expérience du développement de vérifications de santé automatisées, de surveillance synthétique, de cartographies de dépendances et de runbooks opérationnels.

  • Expérience avec Kubernetes, Docker, les services serverless et les environnements cloud natifs.

  • Familiarité avec Apache Kafka et les plateformes de messagerie distribuée.

  • Connaissance des outils de gouvernance et de sécurité cloud ainsi que des solutions CSPM telles que Wiz, Prisma ou CloudGuard.

  • Expérience du suivi opérationnel de services d’IA dans le cloud (Azure AI, AWS Bedrock, Google Vertex AI).

  • Expérience du support d’environnements Linux et Windows via le scripting, l’automatisation, la surveillance et le dépannage.

Pourquoi nous rejoindre ?

  • Une communauté internationale réunissant plus de 110 nationalités différentes.

  • Un environnement où la confiance est au cœur de nos valeurs : 70 % de nos dirigeants ont commencé leur carrière à un poste de niveau débutant.

  • Un système de formation solide grâce à notre Académie interne et à plus de 250 modules de formation disponibles.

  • Un environnement de travail dynamique qui se retrouve régulièrement lors d’événements internes (afterworks, team buildings, etc.).

Amaris Consulting promeut l’égalité des chances. Nous nous engageons à rassembler des personnes issues de parcours divers et à créer un environnement de travail inclusif. À ce titre, nous accueillons les candidatures de toutes les personnes qualifiées, sans distinction de sexe, d’orientation sexuelle, de race, d’origine ethnique, de croyances, d’âge, de situation matrimoniale, de handicap ou de toute autre caractéristique.

ENGLISH VERSION

We are looking for an experienced Site Reliability Engineer (SRE) to support cybersecurity, data risk, and resilience platforms by ensuring their reliability, availability, performance, and operational visibility. The successful candidate will play a key role in maintaining production stability, enhancing observability, supporting incident response, and delivering dashboards and operational insights for engineering, operations, risk, and executive stakeholders.

Location & Work Mode

  • Location: Remote across Canada

  • Work Mode: Fully Remote

Key Responsibilities

  • Maintain and improve the reliability, availability, scalability, and performance of cybersecurity platforms and supporting infrastructure.

  • Monitor system health, identify operational risks, respond to incidents, and drive timely resolution of service-impacting issues.

  • Instrument infrastructure, applications, APIs, data pipelines, databases, and cloud services to provide end-to-end operational visibility.

  • Design, implement, and continuously improve monitoring, alerting, logging, tracing, and observability capabilities across distributed and cloud-native environments.

  • Develop actionable alerts that reduce noise and improve signal quality for faster issue detection and response.

  • Define and track reliability metrics, including SLIs, SLOs, SLAs, error budgets, latency, throughput, error rates, and saturation metrics.

  • Build and maintain dashboards for engineering, operations, product, risk, and executive stakeholders.

  • Continuously improve executive dashboards to support leadership reviews of service health, reliability trends, incidents, risks, and operational performance.

  • Collaborate with engineering, infrastructure, cloud, and cybersecurity teams to identify reliability gaps and implement long-term improvements.

  • Participate in incident response, root-cause analysis, post-incident reviews, and problem management activities.

  • Automate operational tasks, health checks, deployment validation, reporting, and recovery procedures.

  • Support CI/CD, DevOps, and release management processes by validating operational readiness, monitoring coverage, rollback plans, and production support requirements.

  • Contribute to resiliency engineering initiatives, including capacity planning, performance tuning, failover validation, disaster recovery readiness, and resilience testing.

  • Ensure monitoring, alerting, dashboards, and operational processes align with enterprise security, risk, compliance, and governance standards.

Required Qualifications

  • 10+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Infrastructure Engineering, Software Engineering, or Production Operations.

  • Strong experience supporting highly available, distributed, cloud-based, or mission-critical platforms.

  • Hands-on expertise with monitoring, alerting, logging, metrics, tracing, dashboards, and observability practices.

  • Experience instrumenting applications, services, APIs, databases, infrastructure, and cloud components.

  • Strong understanding of SLIs, SLOs, SLAs, error budgets, incident management, capacity management, and operational readiness.

  • Proven experience designing actionable alerts and supporting rapid issue detection and resolution.

  • Experience building operational dashboards for technical teams and executive stakeholders.

  • Strong scripting or programming skills in Python, Bash, PowerShell, Java, or similar languages.

  • Experience with AWS, Azure, or GCP.

  • Experience with Infrastructure as Code (Terraform or equivalent).

  • Familiarity with CI/CD pipelines, DevOps workflows, release management, and production support models.

  • Strong troubleshooting skills across distributed systems, REST APIs, messaging platforms, and service integrations.

  • Familiarity with PostgreSQL, MSSQL, MongoDB, or other relational and non-relational databases.

  • Excellent written and verbal communication skills, including the ability to communicate technical issues to non-technical and executive audiences.

Why choose us

  • An international community bringing together more than 110 different nationalities
  • An environment where trust is central: 70% of our leaders started their careers at the entry level
  • A strong training system with our internal Academy and more than 250 modules available
  • A dynamic work environment that frequently comes together for internal events (afterworks, team buildings, etc.)

Amaris Consulting promotes equal opportunities. We are committed to bringing together people from diverse backgrounds and creating an inclusive work environment. In this regard, we welcome applications from all qualified individuals, regardless of sex, sexual orientation, race, ethnicity, beliefs, age, marital status, disability, or other characteristics.

Create a job alert for this search

Ingénieur(e) Site Reliability (SRE) • Toronto, Canada

Similar jobs

Site Reliability Engineer

CapgeminiToronto, ON, CA
Full-time

Talent Acquisition Business Partner – Strategic Business Unit at Capgemini America Inc.Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d ... Show more

 • Promoted

Site Reliability Engineer

MongoDBToronto, Canada
Full-time

MongoDB's Storage Layer Services (SLS) team is re-architecting the MongoDB cloud storage layer and sits at the heart of our next-generation cloud storage architecture.This relatively new team is bu... Show more

 • Promoted

Manager, Site Reliability Engineering

MastercardToronto, Canada
Full-time

Our PurposeMastercard powers economies and empowers people in 200+ countries and territories worldwide.Together with our customers, we're helping build a sustainable economy where everyone can pros... Show more

 • Promoted

Senior Site Reliability Engineer

MorningstarToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Senior Staff Site Reliability Engineer

CerebrasToronto, ON, CA
Full-time

Become a Staff Site Reliability Engineer at Cerebras Systems, revolutionizing AI inference service reliability.Design innovative solutions for operational challenges.This position is crucial for en... Show more

 • Promoted

Senior Site Reliability Engineer

Guidewire SoftwareToronto, ON, CA
Full-time

At Guidewire, we make software that offers Property and Casualty (P&C) Insurance companies the tools to take care of their customers when they need it the most, whether that’s a time of crisis, a n... Show more

 • Promoted

Site Reliability Engineer

DexianToronto, Ontario, Canada
Full-time

Working Location: Toronto, ON (Hybrid 2 days a week in office).The DevOps and Automation is looking for a Site Reliability Engineer with strong expertise in Dynatrace to ensure the reliability, per... Show more

 • Promoted

Site Reliability Engineer

KyndrylToronto, ON, CA
Full-time +1

Join to apply for the Site Reliability Engineer role at Kyndryl.Direct message the job poster from Kyndryl.Recruitment & Strategic Staffing @Kyndryl | Partnering with IT Consultants in Financial Se... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificToronto, ON, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Remote Senior Site Reliability Engineer Role

ViafouraToronto, ON, CA
Remote
Full-time

Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure.This remote role positions you to improve our platform's performance and sca... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability And Performance

RootlyToronto, Canada
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you'... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalToronto, ON, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Impactful Site Reliability Engineer Fostering Reliability and Performance

RootlyToronto, ON, CA
Full-time

Join as an impactful Site Reliability Engineer, shaping the technical future and enhancing system reliability.Tackle rewarding challenges in a collaborative startup atmosphere.As a key player, you’... Show more

 • Promoted

Site Reliability Engineer

Future Secure AIToronto, ON, CA
Full-time

At Future Secure AI, we're building something genuinely new — and we're looking for people bold enough to build it with us.We work at the frontier of AI, tackling big, real-world problems for globa... Show more

 • Promoted

Senior Site Reliability Engineer

Morningstar Credit Ratings, LLCToronto, ON, CA
Full-time

Investment Services is Morningstar’s internal product group focused on building and maintaining the platforms that power our global data operations.We enable the Managed Investment Data (MID), Refe... Show more

 • Promoted

Site Reliability Engineer (Senior Or Staff), Deployments

AlleyCorpToronto, Canada
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Site Reliability Engineer (Senior or Staff), Deployments

AlleyCorpToronto, ON, CA
Full-time

Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the broader engineering organization.Among these ... Show more

 • Promoted

Site Reliability Engineer (SRE)

Tangerine BankToronto, ON, CA
Permanent

Press Tab to Move to Skip to Content Link.Select how often (in days) to receive an alert:.Tangerine is Canada’s leading direct bank.We offer flexible and accessible banking options, innovative prod... Show more

 • Promoted

Senior SRE Leader: Scale Reliability & Observability

RootlyToronto, Ontario, Canada
Full-time

A fast-growing tech startup in Toronto is seeking an experienced Site Reliability Engineer.The role involves enhancing service performance, owning CI/CD pipelines, and building automation tools.Ide... Show more

 • Promoted

Senior Site Reliability Engineer

iManageToronto, ON, CA
Full-time

SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe.We organize ourselves into distributed teams – SRE teams are anchored t... Show more