Site Reliability Manager, Data Center Networking, Sre (Alton)

Site Reliability Manager, Data Center Networking, Sre (Alton)

25 Sep
|
Google Canada
|
Alton

25 Sep

Google Canada

Alton

Google's Site Reliability Engineering (SRE) team in Waterloo, Ontario is looking for a Site Reliability Manager, Data Center Networking to lead a high-performing team at the intersection of software engineering and large-scale infrastructure. This is a senior leadership role within Google's broader SRE organization, where your work directly shapes the reliability and performance of systems that serve millions of GCP customers worldwide.In this position, you'll be responsible for building a mission-first culture, scaling your leadership through trusted Tech Leads and domain experts, and driving meaningful improvements to incident detection and mitigation across Software-Defined Networking (SDN) infrastructure. It's a role that demands both deep technical expertise and the people leadership skills to move a complex, distributed organization forward.About the Role: Site Reliability Manager, Data Center NetworkingAs the Site Reliability Manager for Data Center Networking, you'll serve as the ultimate execution owner for your team's strategic initiatives. Your primary focus will be on dramatically improving Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM) for incidents — leveraging advanced signaling, tooling, and the integration of new signals into auto-mitigation systems. You'll also set the standard for developer excellence across the SDN ecosystem, influencing the design and rollout of recent network products to ensure they're introduced safely and deliver high reliability.Collaboration is central to this role. You'll partner closely with the PLANET team and sibling SRE shards to define and monitor Network Service Level Objectives (SLOs), co-own blameless postmortems,



and establish end-to-end repair coverage for network infrastructure. Google's SRE culture is grounded in intellectual curiosity, psychological safety, and a commitment to eliminating toil through automation and systems thinking.Benefits and SalaryThis role offers a salary range of CAD $216,000 – $221,000 per year, plus a 20% bonus target, equity, and a comprehensive benefits package. Individual compensation is determined by job-related skills, experience, and relevant education or training. For full details on Google's benefits, visit their official careers page.Job DetailsCompany: GoogleLocation: Waterloo, ON, CanadaRequisition ID: 92399333910946502Pay: CAD $216,000 – $221,000 per year + 20% bonus target + equity + benefitsResponsibilitiesThis role spans strategic leadership, technical influence, and operational ownership. You'll be expected to drive reliability improvements at scale while mentoring and empowering the people around you — all within a culture that values blameless learning and continuous improvement.Build and sustain a cohesive, mission-first culture across multiple locations by scaling leadership through trusted Tech Leads and domain expertsActively prioritize the team's workload to ensure sustained high performance and healthy on-call rotationsOwn execution of the team's strategic efforts,



with a focus on drastically improving MTTD and MTTM for incidents through advanced signaling and toolingIntegrate new signals into auto-mitigation systems to reduce incident impact and manual interventionSet the bar for developer excellence across the SDN ecosystem, influencing the design and safe rollout of new network products (NPIs)Partner with PLANET and sibling SRE shards to define and monitor Network SLOs and establish end-to-end repair coverage for network infrastructureCo-own blameless postmortems to drive systemic improvements following incidentsRequirements / SkillsGoogle is looking for a leader who brings both deep software engineering expertise and proven experience managing technical teams in complex, distributed environments. The ideal candidate thrives in ambiguity, leads with curiosity, and has a track record of improving system reliability at scale.Bachelor's degree in Computer Science or a related technical field, or equivalent practical experience8 years of experience in software development, including work with data structures and algorithms3 years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systemsMaster's degree or PhD in Computer Science, Engineering, or a related field is preferredBachelor's degree in Computer Science or a related technical field or equivalent practical experience. 8 years of experience in software development, and with data structures and algorithms. 3 years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems. Master's degree or PhD in Computer Science or Engineering, or a related field (preferred).

📌 Site Reliability Manager, Data Center Networking, Sre (Alton)
🏢 Google Canada
📍 Alton

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability manager, data center networking, sre (alton) / alton

Subscribe to this job alert:

Get the latest job offers by email for: site reliability manager, data center networking, sre (alton) / alton