Talent.com
Tyk Technologies
Site Reliability EngineerTyk Technologies • CA
Site Reliability Engineer

Site Reliability Engineer

Tyk Technologies • CA
30+ days ago
Job type
  • Full-time
  • Remote
  • Quick Apply
Job description

Who are Tyk, and what do we do?

The Tyk API Management platform is helping to drive the connected world and power new products and services. We’re changing the way that organisations connect any number of their systems and services.Whether internal, external, public or highly encrypted systems, Tyk helps businesses drive value across the retail, finance, telecoms, healthcare, or media industries (to name just a few!)

If you’ve banked online, used an app to check the news, or perhaps even driven a connected car, API’s, and by extension, Tyk, make that possible. Founded in 2015 with offices in London – UK, London – Ontario, Atlanta and Singapore, we have many thousands of users of our B2B platform across the globe. Brands using Tyk range from Lotte, Bell, T Mobile, to RBS, Capital One and Vinci. We have a varied user base hailing from every continent – even Antarctica.

Our Mission

Tyk is on a mission to connect every system in the world. We’ve started by building an API Management platform.

Total flexibility, default remote, radical responsibility

We offer unlimited paid holidays and remote working from anywhere in the world, for everyone, Why? Tyk was founded on the principle of offering flexibility and autonomy to our employees, we believe this allows our employees to achieve their best results. It also means we can build the best possible team, location and working hours are no barrier.

If this sounds like an environment that you believe could work for you then read on to find out more.

The role:

We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform. You will be curious by nature, always looking for ways to improve, as we will look to you for new ideas, solutions and metrics on how we can improve the platform. You will also be our first line of incident management to our clients and will help define our response going forward. This is a great opportunity to become an integral part of Tyk as we continue on our journey.

As a remote first company, you will have the opportunity to work with an industry leading distributed team. Having access to expertise from across the globe will give you both the support and opportunity to help shape not only Tyk’s Cloud platform but also the Tyk as a whole as we continue to grow.

Requirements

Here’s what you’ll be responsible for:

  • Maintaining global Tyk Cloud within SL(A/I/O)s you will help to define
  • Identifying reliability issues and working together with your squad to solve them
  • Identifying and introducing new metrics and building relevant dashboards
  • Participating in the on-call rotation
  • Working with your squad to expand multi-region and multi-cloud reach of the platform
  • Documenting operational knowledge
  • Conducting post-incident analysis
  • Automating common tasks
  • Be a key shaper and contributor to our continuous improvement agenda – be it the clarity of our user stories, how we estimate, communicate with other teams or customers – we expect this role to be advocate of continuous improvement
  • Reliability of our new global Tyk Cloud platform
  • Automation of operations and support
  • Writing and maintaining documentation on SRE processes and policies
  • Recommending and implementing ways of driving operational efficiency and driving down our cost to run, without impacting service
  • Assisting in penetration testing for Cloud through liaising with our provider, providing technical details, and environment setup
  • Incident management

Here’s what we’re looking for:

Experience

  • Strong collaboration skills
  • Launching and operating production scale kubernetes clusters
  • Designing and operating infrastructure on AWS and other providers
  • Operating MongoDB (or other document database) clusters
  • Operating Redis (or other key-value storage) clusters
  • Administering Linux servers
  • Maintaining distributed software
  • Operating Prometheus and Grafana
  • Operating logging collection and analysis systems
  • Participating in the on-call rotation(16:00pm – 4:00am UTC)

Skills:

  • Kubernetes & containers (advanced)
  • AWS / EKS (advanced)
  • Linux (advanced)
  • Terraform and IaC in general (proficient)
  • Helm (proficient)
  • Go (familiar)
  • MongoDB (or similar)
  • Redis (or similar)
  • Monitoring – prometheus, grafana, thanos (familiar)
  • Grasp of networking concepts (subnets, routing, peering, load balancing, NAT, etc.)
  • Common networking protocols (DNS, TCP/IP, HTTP, TLS, UDP)
  • Proactive, energetic, innovative and change oriented

Nice to have:

  • GCP or Azure
  • Bare metal infrastructure engineering
  • API management experience
  • Large scale distributed storage management
  • Familiarity with Rancher
  • CKA/CKAD/CKS
  • Creating and delivering production software in Go language

Benefits

Here’s why you should join us:

  • Everyone has unlimited paid holiday.
  • We have total flexibility in hours, as we believe creativity flows better when our people are given freedom to decide when they are most productive. Everyone is unique after all.
  • Employee share scheme
  • Generous maternity and paternity leave
  • Company retreats

We all share the same vision – we value authenticity, respect, responsibility, independence, honesty, diversity and inclusion and most importantly treating others how you wish to be treated. We look for like-minded people who bring their personalities to work everyday, strive to achieve their personal goals and who are willing to challenge the way we do things, why? – to make what we do even better!

Our values tell the story of Tyk – here’s how:

  • It’s ok to screw up!

We’ve found that it’s often the ‘stupid’ or unexpected ideas that turn out to be the successful ones – so try it, at least we can say we have!

  • The only stupid idea, is the untested one!

It’s in our DNA – starting a business with founders 12 hours apart, giving our gateway away for free – sure, we did that, and we’d do it again!

  • Trust starts with you – make it count!

Trust is a two-way street – instill it from day one!

  • Assume best intent!

We have each other’s back – we’re all on the same team. Think before you speak or act.

  • Make things, better!

Always try to leave things better than when you found them – change is constant, inevitable and embraced! Be that change we want to see.

What’s it like to work here?! check it out: https://tyk.io/worklife/

Tyk is an equal opportunities employer and we are determined to ensure that no applicant or employee receives less favourable treatment on the grounds of gender, age, disability, religion, belief, sexual orientation, marital status, or race, or is disadvantaged by conditions or requirements which cannot be shown to be justifiable.

You can see more about us here https://tyk.io

Create a job alert for this search

Site Reliability Engineer • CA

Similar jobs

Site Reliability Engineer

TELUS DigitalCA
Full-time

Welcome to TELUS Digital — where innovation drives impact at a global scale.As an award-winning digital product consultancy and the digital division of TELUS, one of Canada’s largest telecommunicat... Show more

 • Promoted

Senior Site Reliability Engineer

AlleyCorp, , canada, Canada
Full-time

SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries.Founded in 2013 by security and risk experts Dr.Alex Ya... Show more

 • Promoted

AI Site Reliability Engineer Role

Phizenix, , canada, Canada
Full-time

Join the forefront of AI transformation as our Site Reliability Engineer, focusing on automation and reliability across diverse infrastructure domains.Collaborate with engineering teams to tackle c... Show more

 • Promoted

Site Reliability Engineering — AI Accelerator Infrastructure

PhizenixCA
Full-time

AI to power the transformation of technology.We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible.We value humility and believe in direct communic... Show more

 • Promoted

Intermediate Site Reliability Engineer Role

GitLab, , canada, Canada
Full-time

Advance your site reliability engineering career with GitLab, specializing in Environment Automation.Contribute to ensuring optimal operation and scalability of GitLab environments remotely.As a Si... Show more

 • Promoted

Raptor Engine Site Reliability Engineer

SpaceX, , canada, Canada
Full-time

Join SpaceX as a Raptor Engine Site Reliability Engineer, tackling challenging systems engineering issues.Focus on server management, networking, and High Performance Computing.In this role, you wi... Show more

 • Promoted

Site Reliability Engineer (OpenShift & Infrastructure)

Accion LabsCA
Full-time

Site Reliability Engineer (OpenShift & Infrastructure).Get AI-powered advice on this job and more exclusive features.Direct message the job poster from Accion Labs.Install, configure, upgrade, and ... Show more

 • Promoted

Site Reliability Engineer Role at Illumio

Illumio, , canada, Canada
Full-time

Transform cybersecurity as a Site Reliability Engineer at Illumio.Focus on AWS and Azure infrastructures while based in Sunnyvale, CA, in a vibrant, innovative culture.Illumio is searching for a mo... Show more

 • Promoted

Hadrian Site Reliability Engineer Position

Hadrian Automation, , canada, Canada
Full-time

Join Hadrian as a Site Reliability Engineer, where your mission is to empower end-users with automated solutions and robust device management.Leverage your scripting techniques to enhance workplace... Show more

 • Promoted

Senior Site Reliability Engineer — Kubernetes, AWS & Observability

Thinkific, , canada, Canada
Full-time

A leading e-learning provider in Canada is seeking a Senior Site Reliability Engineer to enhance and secure their infrastructure supporting online course creators.This role involves improving perfo... Show more

 • Promoted

Sr. Site Reliability Engineer

Illumio, , canada, Canada
Full-time

Illumio is the leader in ransomware and breach containment, redefining how organizations contain cyberattacks and enable operational resilience.Powered by the Illumio AI Security Graph, our breach ... Show more

 • Promoted

Senior Site Reliability Engineer

Sage Recruiting Inc., , canada, Canada
Full-time

This range is provided by Sage Recruiting Inc.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Senior Site Reliability Engineer (Founding Role).T... Show more

 • Promoted

Intermediate Site Reliability Engineer, Environment Automation

GitLab, , canada, Canada
Full-time

GitLab is the intelligent orchestration platform for DevSecOps.GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, ... Show more

 • Promoted

Site Reliability Engineer, Client Platform

Hadrian Automation, , canada, Canada
Permanent

Hadrian - Manufacturing the Future.Hadrian is building autonomous factories that help aerospace and defense companies manufacture rockets, satellites, jets, and ships up to 10x faster and up to 2x ... Show more

 • Promoted

Senior Site Reliability Engineer

Thinkific, , canada, Canada
Full-time

Senior Site Reliability Engineer.Senior Site Reliability Engineer.Are you an experienced Site Reliability Engineer looking for a new challenge?.Senior Site Reliability Engineer.Senior Site Reliabil... Show more

 • Promoted

Senior Site Reliability Engineer- Remote

ClickHouse, , canada, Canada
Remote
Full-time

Senior Site Reliability Engineer- Remote.Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most innovative and fast-growing private cloud companies.With more than 3,000 custome... Show more

 • Promoted

Senior Site Reliability Engineer

ShippoCA
Full-time

At Shippo, our vision is bold and clear:.We’re building the backbone of global e-commerce — connecting merchants to carriers worldwide through a single API and intuitive dashboard.We invest in mode... Show more

 • Promoted

DevOps / Site Reliability Engineer

General Matter, , canada, Canada
Full-time

General Matter is enriching uranium in America.Our mission is to restore our country’s ability to make nuclear fuel.Our fuel will help power AI, manufacturing, and other critical industries.It will... Show more

 • Promoted

Site Reliability Engineer (Raptor)

SpaceX, , canada, Canada
Permanent

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not.Today SpaceX is actively developing the technolo... Show more

 • Promoted

Senior Site Reliability Engineer (Remote-First)

VySystemsCA
Remote
Full-time

A leading technology company is seeking a Senior Site Reliability Engineer with robust Kubernetes knowledge to work remotely.Ideal candidates have over 6 years of experience in IT disciplines, prof... Show more