AboutAlayaCare
AtAlayaCareweremore than just a fast-growing SaaS companywerea team of people passionate about transforming home healthcare. Our cloud-based platform empowers care providers around the world to deliver better outcomes for their clients.
With 550 employees across Canada the US Australia and Brazilwereunited by a shared mission and a strong culture of transparency growth and human connection. Whetheryoureearly in your career or a seasoned expertAlayaCareoffers the opportunity to grow your impact your skills and your career.
About the Role
We are seeking a Senior Site Reliability Specialistto join our SRE team. Reporting to the Engineering Manager you is responsible for scaling AWS cloud infrastructure evolving Kubernetes deployment pipelines improving monitoring alerting and resiliency and developing tooling that enables product teams to deliver safely and efficiently. For acquired Azure-based products the focus is on monitoring alert triage and runbook-driven incident response rather than greenfield platform design.
This role owns shared platform services across cloud regions including databases messaging logging search and tenant provisioning. The Senior SRE is expected to lead major infrastructure initiatives and proof-of-concept efforts contribute to technical planning and prioritization as well as partners with the Product teams to reduce operational incidents.
The SRE team also develops and operates AI-driven tools to streamline runbooks accelerate incident response and generate operational insights from platform telemetry to improve reliability and reduce manual effort.
What Youll Do
Development Automation and Tooling
- Design build and maintain infrastructure and platform services including Kubernetes and observability tooling.
- Implement infrastructure as code configuration management and automated testing to ensure reliable repeatable environments.
- Contribute to code and configuration reviews to improve scalabilitymaintainability and reuse.
Reliability and Operations
- Monitor production systems troubleshoot issues and improve logging monitoring alerting and runbooks.
- Participate in on-call and help desk rotations incident response and post-incident reviews to improve long-term reliability.
Requirements and Collaboration
- Partner with Product Engineering and development teams to translate requirements into reliable and operable infrastructure solutions.
- Identify risks across operability security performance and cost and recommend practical trade-offs.
Continuous Improvement
- Contribute to operational quality through runbooks security hardening performance tuning and process improvements.
- Stay current with emerging SRE practices including AI-assisted operations and modern AWS platform patterns.
What You Bring to the Team
- Bachelors or advanced degree in computer science computer engineering or related practical fields with demonstrated experience.
- 5 years of hands-on experience
- Solid hands-on experience with AWS in a multi-account multi-region environment: EKS AWS Organizations IAM and KMS.
- Strong proficiency with Terraform and Infrastructure as Code workflows including Atlantis/GitOps state management and module/provider upgrades.
- Practical experience running workloads on Docker and Kubernetes in production including Gateway API ingress patterns and cluster lifecycle management (upgrades addons node provisioning).
- Strong experience with Linux systems administration and production troubleshooting.
- Proficiency in at least one development or scripting language such as Python Go or Bash.
- Experience with an observability platform (e.g. New Relic OpenSearch CloudWatch OpenTelemetry) and event-driven alerting (e.g. EventBridge SNS PagerDuty).
- Knowledge of system and network security fundamentals including WAF least-privilege IAM secrets management and backup/disaster recovery.
- Experience participating in incident management (on-call triage remediation post-incident review) and writing operational runbooks.
- Hands-on experience operating Aurora MySQL and PostgreSQL in production including migrations performance tuning and backup/restore.
- Strong communication and collaboration skills with the ability to work effectively across technical and non-technical teams in a distributed environment.
- Experience with cloud cost optimization (rightsizing reserved capacity and cost allocation tagging) is an asset.
- Experience with Flux/ArgoCD Karpenter or Ray/Anyscale GPUinfrastructureis an asset.
- Experience monitoring production systems on Azureis an asset.
- Relevant cloud or Kubernetes certifications are an asset.
- Bilingual in French and English is an asset
Why JoinAlayaCare
Work With Purpose
AtAlayaCareyoullhelp build technology that empowers care providers and improves outcomes for patients and families. Every line of code and every customer interactioncontributesto making care more connected accessible and human.
Grow in a High-Trust Culture
We believe in transparency feedback and assuming positive intent. Hereyoullfeel safe to share your ideas and career goals and be supported to achieve them through mentorship career mobility and a promote-from-within philosophy.
Balance That Works for You
We value flexibility and well-being. From Wellness Fridays to volunteer time off to flexible vacation we make sure you have the space to recharge contribute to your community and live your best life.
BenefitsThat Matter
- Equity in a well-funded scaling company.
- Comprehensive health benefits telemedicine and lifestyle spending accounts.
- Parental leave top-up and family support programs.
Inclusive by Design
We celebrate diverse perspectives and foster belonging through our DEIB initiatives. Employee-led events summits and social activities both in-person and virtual create meaningful connections across our global teams.
Location and Work Model
This role is based in Montreal. At AlayaCare our hybrid model includes 2 set in-office collaboration days/week and it is expected that team members are present in the office on those days to foster connection innovation and teamwork.
Ready to Join Us
Apply today and be part of a company that makes a real difference in the future of home and community care. Not the right role for you Share thispostingwith someone who might be a great fit.
AlayaCareuses AI tools during our hiring process to support fair consistent and objective decision-making. Someinitialscreening steps may be automated to helpidentifyqualified candidates. If your application is declined automatically you may request a human review.
Werecommitted to creating a workplace where everyone belongs. If you require accommodation during the application process please reach out to.
Required Experience:
Senior IC
Employment Type : Full Time
Experience: years
Vacancy: 1