Talent.com
Xsolla
Site Reliability Engineer (Monetization)Xsolla • Montreal, Quebec, Canada
Site Reliability Engineer (Monetization)

Site Reliability Engineer (Monetization)

Xsolla • Montreal, Quebec, Canada
19 days ago
Salary
CA$120,000.00 yearly
Job type
  • Full-time
Job description

ABOUT YOU

We are looking for a Site Reliability Engineer (Monetization) who is pragmatic product-minded and equally comfortable writing code and running production systems to join our Infrastructure departments SRE team. The best candidate will be someone who thrives in a fast-paced highly collaborative and exceptionally dynamic setting and is excited to own the application-level infrastructure and reliability of a high-traffic commerce domain end to end - from deploy pipelines and Kubernetes manifests to SLOs capacity planning and production readiness.

Strong Kubernetes observability and software engineering skills are essential along with experience in operating production services in a cloud environment (GCP/GKE or comparable) and partnering closely with product development teams. The ability to hold a dual perspective - understanding both how developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions early will be key to your success in this role.

This is a hybrid embedded role: you remain part of the SRE organization (practices standards duty rotation) while being functionally embedded into the Monetization product domain. Youll build long-term working relationships with the domains engineering teams own a meaningful share of their application infrastructure execution and co-author the reliability practices used company-wide.

If youre passionate about making complex distributed systems boringly reliable and love building the commerce and monetization backbone that lets game developers around the world get paid we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA companies partner with Xsolla to help them fund distribute market and monetize their games. Grounded in the belief in the future of video games Xsolla is resolute in the mission to bring opportunities together and continually make new resources available to creators. Headquartered and incorporated in Los Angeles California Xsolla operates as the merchant of record and has helped over 1500 game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win developers have all the things needed to enjoy the game.

For more information visit .

Responsibilities

  • Own the application-level infrastructure of the Monetization domain: Helm charts Terraform configurations Kubernetes deployments runtime configuration and service-level networking and integrations
  • Own the domains observability: design and implement SLOs/SLIs monitors alerts and dashboards for critical services on Datadog and OpenTelemetry-based tooling
  • Help to set up and evolve CI/CD pipelines for domain services (GitLab CI GitHub Actions) including deploy and rollback automation
  • Perform capacity planning and performance tuning ahead of expected load - product launches sales events and regional rollouts - including load testing and performance regression investigation
  • Run Production Readiness Reviews for new services and major changes; define and enforce what production-ready means for the domain
  • Support domain incident response: assist with deep investigation of complex incidents contribute to post-mortems drive follow-up reliability improvements and maintain runbooks
  • Build domain-specific automation that reduces operational toil: runbook automation deploy helpers recurring operational scripts
  • Maintain and drive a forward-looking reliability roadmap for the domain together with product engineering leads
  • Participate in product team planning refinements and architecture reviews bringing the reliability perspective before design decisions become expensive to change
  • Co-author company-wide SLO/SLI capacity and operational standards together with the broader SRE team; contribute improvements directly to shared SRE-operated subsystems
  • Participate in the SRE duty rotation supporting developers across the company

Qualifications & Skills

  • 3 years of proven SRE DevOps or platform engineering experience: on-call or incident response duty SLO/monitoring ownership deploy pipeline and infrastructure work for production services
  • Software development background: you have built and shipped backend services not only operated them - comfortable reading application code during an investigation and writing production-quality automation in at least one language (e.g. Go PHP)
  • Hands-on Kubernetes experience:Helm manifests deploy strategies debugging application-level performance and networking issues (GKE or another managed Kubernetes)
  • Solid observability practice: building monitors dashboards and SLOs/SLIs on a modern platform (Datadog preferred; Prometheus/Grafana experience also relevant) familiarity with OpenTelemetry
  • Infrastructure as Code exposure (Terraform/Terragrunt) for collaboration with platform teams
  • GCP experience (IAM networking managed services)
  • Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions)
  • Programming/scripting proficiency sufficient to build automation and tooling (e.g. Python Go or Bash)
  • Practical experience with incident response post-mortems and driving reliability improvements from incidents
  • Strong collaboration and communication skills this role works embedded with product development teams daily
  • Experience in payments fintech e-commerce or gaming high-traffic transactional systems

Nice to Have:

  • Kubernetes certifications
  • Google Cloud Platform certifications
  • HashiCorp certifications
$120000 - $160000 a yearSalary varies depending on experience level and location.

Benefits

We are passionate about fostering a supportive environment for our team so we prioritize the physical mental and emotional well-being of our employees and their families through a comprehensive Benefits Program. This includes medical dental and vision PTO and a personalized career roadmap for each employee. By investing in professional development through training and educational opportunities we ensure that our team thrives both personally and professionally. Together were not just building a business; were cultivating a community that values creativity collaboration and the transformative power of play.

Equal Employment Opportunity Statement

Xsolla is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race color religion sex national origin age disability sexual orientation gender identity or any other characteristic protected by law. We consider qualified applicants with criminal histories in accordance with the Fair Chance Act.

Criminal History Consideration

For the Site Reliability Engineer (Monetization) position we will conduct a background check that may include the following:

  • Criminal history check
  • Employment verification
  • Education verification

Relevance to Job Responsibilities

The background check is relevant to this position because of the following role responsibilities:

  • Accessing confidential company data
  • Handling infrastructure that processes sensitive financial transactions
  • Ensuring compliance with regulatory requirements

Rights Under the Fair Chance Act

Applicants are encouraged to inquire about their rights under the Fair Chance Act. If you have questions regarding our hiring practices please contact emailprotected.

By submitting the following job application form you consent to Xsolla processing your data for career-related inquiries and potential employment opportunities. We process your data in accordance with this Xsolla Privacy Notice for Job Applicants. Please direct any inquiries regarding your data privacy to emailprotected.

We may use artificial intelligence (AI) tools to support parts of the hiring process such as reviewing applications analyzing resumes or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed please contact us.

Required Experience:

IC


Employment Type : Full-Time
Experience: years
Vacancy: 1
Yearly Salary Salary: 120000 - 160000
Create a job alert for this search

Site Reliability Engineer (Monetization) • Montreal, Quebec, Canada

Similar jobs

Site Reliability Engineer

ApTaskMontréal, Quebec, Canada
Full-time

Direct message the job poster from ApTask Looking for an intermediate between 2 to 5 years' experience.The Application Infrastructure (Al) department is seeking a Site Reliability Engineer (SRE) to... Show more

 • Promoted

Site Reliability Engineer

Vertex Elite LLCRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Duration: ContractKey Skills:Monitoring / Observability tools - Dynatrace, ELK etc.Platform/ cloud Observability - OpenShift, Prometheus / Azure Cloud etc.Key Responsibilities:Collaborate with vari... Show more

 • Promoted

Sr. Engineer

TechDoQuestmontreal (administrative region), qc, Canada
Full-time

Perform icing numerical simulations on complex aerodynamic configurations.Prepare, execute, and analyze high‑lift and icing wind tunnel test campaigns, including CFD and certification.Architect and... Show more

 • Promoted

Senior Full-Stack Engineer - Accessibility & Inclusive Tech (Remote)

Accessibility Partners CanadaMontreal (administrative region), QC, CA
Remote
Full-time

A leader in accessible technology is seeking a Senior Full-Stack Developer to create equitable and accessible digital systems.This role involves developing both front-end and back-end systems, focu... Show more

 • Promoted

Site Reliability Engineer

Hunter BondMontréal, Canada
Full-time

Role: DevOps EngineerClient: Most Elite Tech Firm in CanadaCompensation: Up to $200k CAD + Bonus + PackageLocation: MontrealOverviewAn Elite FinTech Firm is looking for a highly talented DevOps Eng... Show more

 • Promoted

Senior Database Reliability Engineer New Montreal, Canada

AppDirect, Incmontreal (administrative region), qc, Canada
Full-time

Become a digital, global citizen and enable the new generation of digital entrepreneurs around the world.AppDirect offers a subscription commerce platform to sell any product, through any channel, ... Show more

 • Promoted

Director of Engineering — Platform & Reliability (Remote)

CliniaMontreal (administrative region), QC, CA
Remote
Full-time

A tech-driven health company in Canada is seeking a Director of Engineering to lead an engineering team of 25.You will manage delivery, ensure platform reliability, and set engineering standards wh... Show more

 • Promoted

Senior Site Reliability Engineer Focused on Kubernetes Infrastructure

Chainlink LabsMontreal (administrative region), QC, CA
Full-time

Elevate decentralized architecture as a Senior Site Reliability Engineer.Spearhead Kubernetes-based infrastructure for decentralized applications, driving scalability, security, and operational eff... Show more

 • Promoted

Intact Hybrid Site Reliability Engineer

IntactMontreal
Full-time

Join Intact as a Site Reliability Engineer and elevate operational reliability across cloud platforms.This hands-on role employs Azure, AWS, and GCP expertise to manage incidents effectively.The SR... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsMontreal (administrative region), QC, CA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more

 • Promoted

Senior Ii Site Reliability Engineer

Akamai TechnologiesRivière-Des-Prairies-Pointe-Aux-Trembles, Canada
Full-time

Job DescriptionJoin our SRE team! Our team uses large datasets to analyze and measure the performance and reliability of our platform.We are networking data scientists: we combine our knowledge of ... Show more

 • Promoted

Senior Site Reliability Engineer

ThinkificMontreal (administrative region), QC, CA
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Site Reliability Engineer

BasetenMontréal, Canada
Full-time

About BasetenBaseten powers mission‐critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.By uniting applied AI resear... Show more

 • Promoted

SAR Payload Systems Engineering Advisor

Sky Systems, Inc. (SkySys)montreal (administrative region), qc, Canada
Full-time

SAR Payload Systems Engineering Advisor.Duration: 12 months – 40 hours per week.Location: Montreal – West Island, 4 days per week in office.Salary: CAD $125K - $140K Annually with Standard Benefits... Show more

 • Promoted

Site Reliability Engineer

TELUS DigitalMontreal (administrative region), QC, CA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouseMontreal (administrative region), QC, CA
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Site Reliability Engineer

SecurityScorecardmontreal (administrative region), qc, Canada
Full-time

SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries.Founded in 2013 by security and risk experts Dr.Alex Ya... Show more

 • Promoted

Site Reliability Engineer

MaintainXMontreal
Full-time

MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 13,000+ customers including Duracell, Shell, Cintas, and Brenntag.We raised $150M in Series D funding ... Show more

 • Promoted

Site Reliability Engineer - Tech Talent International

Tech Talent InternationalMontreal
Full-time

Join Tech Talent International as a Senior Site Reliability Engineer, specializing in Automation & Observability, located in Montreal.This hybrid role focuses on enhancing production efficiency and... Show more

 • Promoted

Senior Maintenance & Reliability Leader

PharmascienceMontreal (administrative region), QC, CA
Full-time

A leading pharmaceutical company in Montreal is looking for a Maintenance Manager to oversee and improve the maintenance program for manufacturing and packaging equipment.The successful candidate w... Show more