Site Reliability Manager, Data Center Networking, SRE (London)

Site Reliability Manager, Data Center Networking, SRE (London)

23 Sep
|
Google Canada
|
London

23 Sep

Google Canada

London

Google's Site Reliability Engineering (SRE) team in Waterloo, Ontario is looking for a Site Reliability Manager, Data Center Networking to lead a high-performing team at the intersection of software engineering and large-scale infrastructure. This is a senior leadership role within Google's broader SRE organization, where your work directly shapes the reliability and performance of systems that serve millions of GCP customers worldwide.

In this position, you'll be responsible for building a mission-first culture, scaling your leadership through trusted Tech Leads and domain experts, and driving meaningful improvements to incident detection and mitigation across Software-Defined Networking (SDN) infrastructure. It's a role that demands both deep technical expertise and the people leadership skills to move a complex, distributed organization forward.

About the Role: Site Reliability Manager, Data Center Networking

As the Site Reliability Manager for Data Center Networking, you'll serve as the ultimate execution owner for your team's strategic initiatives. Your primary focus will be on dramatically improving Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM) for incidents — leveraging advanced signaling, tooling, and the integration of new signals into auto-mitigation systems. You'll also set the standard for developer excellence across the SDN ecosystem, influencing the design and rollout of new network products to ensure they're introduced safely and deliver high reliability.

Collaboration is central to this role. You'll partner closely with the PLANET team and sibling SRE shards to define and monitor Network Service Level Objectives (SLOs), co-own blameless postmortems, and establish end-to-end repair coverage for network infrastructure.



Google's SRE culture is grounded in intellectual curiosity, psychological safety, and a commitment to eliminating toil through automation and systems thinking.

Benefits and Salary

This role offers a salary range of CAD $216,000 – $221,000 per year, plus a 20% bonus target, equity, and a comprehensive perks package. Individual compensation is determined by job-related skills, experience, and relevant education or training. For full details on Google's benefits, visit their official careers page.

Job Details

Company: Google

Location: Waterloo, ON, Canada

Requisition ID: 92399333910946502

Pay: CAD $216,000 – $221,000 per year + 20% bonus target + equity + benefits

Responsibilities

This role spans strategic leadership, technical influence, and operational ownership. You'll be expected to drive reliability improvements at scale while mentoring and empowering the people around you — all within a culture that values blameless learning and continuous improvement.

- Build and sustain a cohesive, mission-first culture across multiple locations by scaling leadership through trusted Tech Leads and domain experts
- Actively prioritize the team's workload to ensure sustained high performance and healthy on-call rotations
- Own execution of the team's strategic efforts,



with a focus on drastically improving MTTD and MTTM for incidents through advanced signaling and tooling
- Integrate new signals into auto-mitigation systems to reduce incident impact and manual intervention
- Set the bar for developer excellence across the SDN ecosystem, influencing the design and safe rollout of new network products (NPIs)
- Partner with PLANET and sibling SRE shards to define and monitor Network SLOs and establish end-to-end repair coverage for network infrastructure
- Co-own blameless postmortems to drive systemic improvements following incidents

Requirements / Skills

Google is looking for a leader who brings both deep software engineering expertise and proven experience managing technical teams in complex, distributed environments. The ideal candidate thrives in ambiguity, leads with curiosity, and has a track record of improving system reliability at scale.
- Bachelor's degree in Computer Science or a related technical field, or equivalent practical experience
- 8 years of experience in software development, including work with data structures and algorithms
- 3 years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems
- Master's degree or PhD in Computer Science, Engineering, or a related field is preferred

Bachelor's degree in Computer Science or a related technical field or equivalent practical experience. 8 years of experience in software development, and with data structures and algorithms. 3 years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems. Master's degree or PhD in Computer Science or Engineering, or a related field (preferred).

📌 Site Reliability Manager, Data Center Networking, SRE (London)
🏢 Google Canada
📍 London

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability manager, data center networking, sre (london) / london

Subscribe to this job alert:

Get the latest job offers by email for: site reliability manager, data center networking, sre (london) / london