Staff Site Reliability Engineer - C$124,200 - C$166,700 A Year (Toronto)

Staff Site Reliability Engineer - C$124,200 - C$166,700 A Year (Toronto)

01 Aug
|
Walt Disney Animation Studios
|
Toronto

01 Aug

Walt Disney Animation Studios

Toronto

Staff Site Reliability Engineer (Staff SRE)Job ID10135906LocationVancouver, CanadaBusinessWalt Disney Animation StudiosDate postedOct. 31, 2025Job Summary:Walt Disney Animation Studios’ world‑class filmmakers, artists, and technical collaborators create the magic of animation. Bring your unique talents, passion and ideas to our team and prepare to play in a creative, artist‑friendly environment.We are seeking a Staff SRE with expertise in Linux platform systems administration, software development (e.G. Python, Go, Java, Node), CI pipeline tools (e.G. Jenkins), Git source management, cloud hosting (AWS, GCP & Azure), container computing (e.G. Docker, OCI), and web technologies. The ideal candidate will enjoy the diversity and challenges of working at various levels in the foundational deployment stack, from defining configuration management to developing CI/CD infrastructure and processes.This role resides within the Platform and Infrastructure team at Walt Disney Animation Studios (WDAS), and we build the tools and manage the infrastructure that artists use daily to create our celebrated animated content. The SRE team within Platform Engineering is focused on optimizing service deployments and improving the availability, latency, performance, efficiency, and observability of systems at WDAS. All projects have in common pursuit of simple and performant solutions to complex problems using Agile and DevOps methodologies as part of high‑energy, proficient teams.Critical to success in this role is an aptitude for working collaboratively with a technical team. You will help to develop and drive requirements and strategies while also supporting services and core services infrastructure.Our studio thrives from a wide variety of technical backgrounds and experiences, so we encourage applicants to apply even if they have experiences not specified below. Bring your unique talents, passion and ideas to our team, and be a part of Disney’s creative legacy!ResponsibilitiesAs Staff SRE, you will translate ideas into tangible products that shape experiences by focusing on a systematic approach to automation, resiliency, efficiency, stability, security, performance, and capacity management, as well as documentation. You will serve as a subject matter expert in multiple areas and be looked at by your fellow team members as a "go to" individual; you are someone who hasa clear understanding of, and can thoroughly elaborate on SRE principles and best practices to a given audience. To be successful in this role you will continuously uphold and improve all the relevant reliability aspects for our services, with an increased focus on SLIs and SLOs, while raising the reliability of a variety of large‑scale user‑facing and internal services. As Staff SRE, you will maintain a strong understanding of stakeholder workflows and requirements, and then be able to translate the targeted solutions into an end‑to‑end architectural design.You will work with engineering, creative and production teams in an extremely collaborative and high‑energy environment to brainstorm, architect, gather requirements,



troubleshoot, and provide stellar customer support. You are passionate about constantly learning, applying technology to solve complex problems, and are a highly motivated, optimistic, proactive, creative thought leader and project manager.Additional Responsibilities Include:Support a wide range of on‑premises and cloud deployments using infrastructure‑as‑code, self‑healing, and security automation patterns and can facilitate others to use the Infrastructure as Code paradigmDeploy and manage a wide array of on‑premises and cloud deploymentsDevelop useful telemetry, alerts, and response to reduce Mean Time To Repair (MTTR)Collaborate and provide technical excellence within and across teamsConsult on best practices and develop tools to enable smooth adoptions of valuable service reliability practices and methodsIdentify areas of improvement in reliability, efficiency, and operationsBuild tools to help your SRE team quickly pinpoint, isolate and resolve issues related to infrastructure, platform services and applicationsContinuously refine monitoring processes, configurations, and thresholdsPractice and promote sustainable incident response and blameless postmortemsDevelop runbooks and tools to streamline processes and shorten problem resolution timeWrite code that improves scalability, performance, maintainability, and securityAdd, tune and maintain alert configurations and documentation as neededDevelop and improve CI/CD processes to improve release cadence and successUse Chaos Engineering principles and methodologies to test what you build under real‑world conditionsMentor SREs, Sysadmins, and Systems Engineers in technical and non‑technical SRE responsibilitiesRequired EducationBS in Computer Science, Computer Engineering, Electrical Engineering or related fieldKey Qualifications:7+ years of experience in SRE, devops, technical operations, systems engineering, software engineering or related disciplineProficient, collaborative, & experienced in building reliable, scalable, enterprise systemsExcellent communication skills, both verbal and writtenPassionate and curious about ways to leverage technology while continually learningEfficiently skilled with the use of containers and container orchestration systems in enterprise production environments (e.G. Docker, Kubernetes, Rancher, AWS ECS and EKS)Experience with configuration management and infrastructure as code (e.G. Terraform, Helm, Cloud Formation, Ansible, Puppet, and Ansible)Comfortable in one or more of the following languages (Python, Java, Scala, Go, Rust, Ruby, or similar)Skilled in Cloud/PaaS/SaaS Environments (e.G. AWS, Azure, Google Cloud Compute)Hands‑on experience using source control (Git, GitHub)



and feature branching strategiesExperience with continuous integration tools (e.G. Jenkins, Gitlab CI/CD, AWS CodeBuild, CodeDeploy, Spinnaker)Knowledge of best practices and IT operations in an always‑up, always‑available servicePossess expertise in scalable testing, automation, continuous integration frameworks and best practicesExperience in SDLC, distributed systems, networking, hardware, logistics and operations or capacity planningUNIX/Linux administration, troubleshooting, performance tuning, and securityExperience with DevOps methodologies and/or SREExperience with monitoring and observability tooling such as Datadog, Prometheus, and GrafanaExperience with automating infrastructure, deployment and testing using tools like Cloudformation, Ansible or TerraformExperience with Service Level Objectives and Error BudgetsUnderstanding of the principles and methodologies behind Chaos EngineeringBonus Qualifications:Expertise in web server administrationThe Walt Disney Company is an Equal Opportunity Employer.The hiring range for this position in British Columbia, Canada is C$124,200 to C$166,700 CAD per year. The base pay actually offered will take into account internal equity and also may vary depending on the candidate’s geographic region, job‑related knowledge, skills, and experience among other factors. A full range of medical, financial, and/or other variable pay or benefits, may be offered dependent on the level and position offered.About Walt Disney Animation Studios:Combining masterful artistry and storytelling with groundbreaking technology, Walt Disney Animation Studios is a filmmaker‑driven animation studio responsible for creating some of the most beloved films ever made. Disney Animation continues to build on its rich legacy of innovation and creativity, from the first fully‑animated feature film, 1937's “Snow White and the Seven Dwarfs,” to the upcoming fall 2024 feature, “Moana 2.” Among the studio's timeless creations are “Pinocchio,” “Sleeping Beauty,” “The Jungle Book,” “The Little Mermaid,” “The Lion King,” “Frozen,” “Zootopia,” and “Encanto.”About The Walt Disney Company:The Walt Disney Company, together with its subsidiaries and affiliates, is a leading diversified international family entertainment and media enterprise that includes three core business segments: Disney Entertainment, ESPN, and Disney Experiences. From humble beginnings as a cartoon studio in the 1920s to its preeminent name in the entertainment industry today, Disney proudly continues its legacy of creating world‑class stories and experiences for every member of the family. Disney’s stories, characters and experiences reach consumers and guests from every corner of the globe. With operations in more than 40 countries, our employees and cast members work together to create entertainment experiences that are both universally and locally cherished.This position is with Walt Disney Pictures, which is part of a business we call Walt Disney Animation Studios.#J-18808-Ljbffr

📌 Staff Site Reliability Engineer - C$124,200 - C$166,700 A Year (Toronto)
🏢 Walt Disney Animation Studios
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: staff site reliability engineer - c$124,200 - c$166,700 a year (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: staff site reliability engineer - c$124,200 - c$166,700 a year (toronto) / toronto