Intermediate to Senior Staff Site Reliability Engineer (Canada)

Intermediate to Senior Staff Site Reliability Engineer (Canada)

11 Sep
|
GitLab
|
Canada

11 Sep

GitLab

Canada

Who you are

- We don't expect every candidate to have experience with every technology in our environment. We're looking for engineers with strong technical fundamentals, a growth mindset, and the ability to learn quickly

- Experience keeping production systems reliable, combining an operations mindset with real software engineering practice

- Experience building net-current infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch
- The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes

- Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your level

- Hands-on experience with at least one major cloud provider (GCP or AWS)

- Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions

- Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure

- Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment
- A track record of using automation, and increasingly AI, to reduce toil and improve how you and your team work

- Alignment with GitLab's values and a commitment to working in accordance with them

- Please note that we welcome interest from candidates with varying levels of experience; many successful candidates do not meet every single requirement

- Additionally, studies have shown that people from underrepresented groups are less likely to apply to a job unless they meet every single qualification. If you're excited about this role, please apply and allow our recruiters to assess your application

What the job involves





- Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale

- They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure

- This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department

- Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience and our hiring needs

- We hire Site Reliability Engineers from Intermediate through Senior Staff across multiple Infrastructure Platforms teams

- We'll support you in becoming successful with GitLab's tools, systems, and ways of working

- We’ll calibrate your level throughout the interview process based on the scope and impact of your experience :

- Intermediate: You independently deliver meaningful reliability improvements within a defined area

- Senior: You own complex reliability work end to end and raise the effectiveness of your team

- Staff: You shape reliability across multiple teams, solving systemic problems and creating approaches others can reuse

- Senior Staff: You set technical direction across a broader Infrastructure area and influence reliability strategy at organizational scale

- Keep user-facing services and production systems reliable, scalable, and efficient

- Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows

- Operate and troubleshoot production systems on Kubernetes,



including deployments, rollouts, and scaling

- Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps

- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately

- Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages

- Take part in incident response and post-incident reviews, turning learnings into changes in automation and process

- Document runbooks, architecture decisions, and reviews so your findings become repeatable practices

The application process

- Recruiter Screen: A conversation about your background, what you're looking for, and the level and teams that fit, so we can point your process in the right direction

- Core Technical: The shared assessment every SRE candidate takes, regardless of eventual team. A low-stress, collaborative discussion covering source code, system architecture, and incident review

- Peer Technical: Team-specific depth, run by SREs from the team you're most likely to join, focused on the problems that team actually works on

- Hiring Manager Interview: A conversation about ownership, judgment, execution, collaboration, and growth, the non-technical signals that make an SRE effective at GitLab

- Skip-Level Interview: A conversation with a senior leader on values alignment, and how you'll work across teams

- After your interviews, we consider your performance alongside our current hiring needs to confirm the level and team where you'll do your best work. Interview results are a major factor, and final placement also reflects our active hiring priorities at the time

Benefits

- We offer benefits to manage your health, wealth, and well-being regardless of location

- Flexibility in schedule to be there for life’s important moments

- Equity compensation & Employee Stock Purchase Plan offered

- Generous Paid Time Off

📌 Intermediate to Senior Staff Site Reliability Engineer (Canada)
🏢 GitLab
📍 Canada

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: intermediate to senior staff site reliability engineer (canada) / canada

Subscribe to this job alert:

Get the latest job offers by email for: intermediate to senior staff site reliability engineer (canada) / canada