Platform - Senior Manager, Site Reliability Engineering (Toronto)

Platform - Senior Manager, Site Reliability Engineering (Toronto)

06 Aug
|
Tubi - Canada
|
Toronto

06 Aug

Tubi - Canada

Toronto

Senior Manager, Site Reliability Engineering Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience.

We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation. We are seeking an experienced and visionary Senior SRE Manager to lead and grow our newly built Site Reliability Engineering team. You are more than a people manager or a tech lead; You will build and mentor a team of talented engineers, foster a culture of blameless learning and continuous improvement, and champion the engineering practices that allow us to balance rapid innovation with rock-solid stability.

You will be a key influencer in our engineering leadership, partnering with peers across the organization to ensure reliability is a shared responsibility and a core tenet of our engineering culture. Lead, mentor, and grow a team of Site Reliability Engineers. Foster a culture of innovation and technical excellence where engineers feel empowered to do their best work.

Provide personalized coaching, create professional development plans, and guide the careers of senior and emerging talent within the team. Define team rituals - runbook reviews, game days, and incident retros - that reinforce quality and learning.

Strategic Planning & Vision: Define and drive the multi-year technical strategy and vision for Tubi’s observability, and automation platforms. Champion a data-driven approach to reliability, using Service Level Objectives (SLOs) and error budgets to facilitate productive conversations about risk and feature velocity.

Operational Excellence & Incident Management: Own the end-to-end availability, performance, and efficiency of our critical user-facing services. Champion a rigorous, blameless, and data-driven post-mortem culture to ensure we learn from both successes and failures, driving eng teams for systemic fixes and automation to prevent the recurrence of incidents. Streamline and improve our existing processes and practices, and collaborate with other teams to enhance our production release standards by improving current processes.





Define and tune a 24×7 on-call rotation for low noise and rapid response; Financial & Vendor Management: Own the SRE budget, tooling, and headcount. Manage relationships with key third-party vendors for our observability and SRE related AI platforms, work with infra lead and finance team for contract negotiations and ensure we derive maximum value from our investments. Cross-Functional Collaboration: Act as a key influencer and strategic partner to leaders in Software Engineering, Product Management, and Infra/Sec.

The AI Mandate : Building the Future of Observability with AI. You will not just manage a team that uses AI; you will lead the charge in building an AI-native SRE function. Developing and executing the strategy for integrating AIOps and machine learning into our observability stack.

Accelerating Automation with AI: Championing the effective and responsible use of AI-assisted coding tools (e.g., You will set the standards and practices to leverage these tools to accelerate the development of automation, operational tooling, and infrastructure code.

Building the Business Case: Building the techno-economic case for new AI tooling, managing vendor relationships, and ensuring the cost-effective and secure implementation of these powerful systems. You must be able to articulate the ROI of these investments in terms of reduced downtime, improved operational efficiency, and faster incident resolution.

Fostering Critical AI Literacy: Fostering a culture that can critically evaluate, debug, and learn from the outputs of AI systems. This involves extending our blameless post-mortem philosophy to AI-driven actions and recommendations, ensuring that the team remains in control and understands the "why" behind automated decisions. 8+ years of experience in a technical field, with at least a year in an engineering leadership position managing SRE, DevOps, or Production Engineering teams. ~ A deep, principled understanding of SRE tenets, including Service Level Indicators (SLIs), SLOs, error budgets, toil reduction, and capacity planning. ~ Exceptional communication, negotiation, and influencing skills,



with the ability to articulate complex technical concepts and strategies to both technical and non-technical stakeholders at all levels of the organization. ~ A strong technical background as a hands-on software engineer or site reliability engineer prior to moving into management. Deep knowledge of AWS services (especially networking, IAM, EKS, ALBs/NLBs, Route 53, CloudWatch).

Proven experience with Kubernetes in production (EKS preferred), including service exposure, networking, and availability engineering. ~ PagerDuty, FireHydrant), deployment-safety tooling (e.g., Pursuant to local pay disclosure requirements, the pay range for this role, with final offer amount dependent on education, skills, experience, and location is as listed annually below. This role is also eligible for an annual discretionary bonus, long-term incentive plan, and various benefits including medical/dental/vision, insurance, vacation/paid time off and other benefits in accordance with applicable plan documents. 164,600 - $235,100 CAD For all salaried employees, in lieu of the FOX Vacation policy, Tubi offers a Flexible Time Off Policy to manage all personal matters. For all full-time, regular employees, in lieu of FOX Paid Parental Leave, Tubi offers a generous Parental Leave Program, which allows parents twelve (12) weeks of paid bonding leave (top up in Canada) within the first year of birth, adoption, surrogacy, or foster placement of a child in addition to applicable government leave program(s) and FOX’s short-term disability policy (if applicable).

This time is 100% paid through a combination of any applicable government leaves and wage-replacement programs in addition to contributions made by Tubi. For all full-time, regular employees, Tubi offers a monthly wellness reimbursement. Tubi offers the world's largest collection of Hollywood movies and TV shows, thousands of creator-led stories and hundreds of Tubi Originals made for the most passionate fans.

Headquartered in San Francisco and founded in 2014, Tubi is part of Tubi Media Group, a division of Fox Corporation. We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, gender identity, disability, protected veteran status, or any other characteristic protected by law. We will consider for employment qualified applicants with criminal histories consistent with applicable law. #

📌 Platform - Senior Manager, Site Reliability Engineering (Toronto)
🏢 Tubi - Canada
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: platform - senior manager, site reliability engineering (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: platform - senior manager, site reliability engineering (toronto) / toronto