About FableGlobal enterprises work with Fable to make products more accessible for over one billion people who live with disabilities. Our customers include global leaders like Walmart, Slack, and Shopify. Fable was featured on the Forbes Accessibility 100 list in 2025, awarded Rapid Company 's Most Innovative Companiesin Design, and has received accolades from global entities such as the World Summit Awards and the UN‑endorsed Zero Project.About the RoleAs a Senior Site Reliability Engineer at Fable, you will play a critical role in ensuring the reliability, scalability, and efficiency of our platform as we continue to grow. Fable 's products support organizations in building more accessible digital experiences, and the reliability of our infrastructure is essential to delivering that impact. You will work across our platform and product systems to ensure they are stable, performant, and cost‑efficient, while enabling teams to move quickly and safely.As AI‑powered capabilities increasingly become part of modern product experiences, you will also help ensure Fable 's infrastructure is ready tosupport AI workloads—balancing reliability, performance, and cost while enabling teams to safely experiment and scale new capabilities. Reporting to the Director of Technical Operations, this role works closely with teams across Engineering and Product. It is ideal for someone who enjoys hands‑on technical work while taking ownership of system health, tooling, and operational excellence, and who is excited to help shape Fable 's approach to infrastructure, reliability, and platform engineering over time.ResponsibilitiesReliability, Infrastructure & PlatformDesign, build, and maintain reliable, scalable, and secure infrastructure for Fable 's product servicesImprove system observability, monitoring, and alerting to ensure high availability and fast incident responseContribute to and evolve SRE practices, including SLIs/SLOs, incident management, and postmortemsSupport and improve CI/CD pipelines and deployment processesIdentify and reduce operational complexity across systems and toolingWork across infrastructure and application layers to diagnose and resolve reliability and performance issues, including making targeted improvements to application code when neededSupport infrastructure and platform capabilities required for AI/ML‑powered features, including scaling, performance,
and reliability considerationsCost Efficiency & PerformanceMonitor and optimize infrastructure costs across cloud environmentsContribute to capacity planning and cost forecasting for infrastructure and servicesIdentify opportunities to improve performance and efficiency at the system levelEvaluate and optimize the cost and performance of compute‑intensive workloads (e.G., AI/ML services), ensuring efficient resource usage and scalabilityVendor & Tooling OwnershipWork with third‑party vendors and tools that support Fable 's infrastructure and operationsHelp evaluate, select, and manage tools and services to support platform reliability and scalabilitySupport vendor‑related troubleshooting and ongoing service improvementsCross‑functional CollaborationPartner with Engineering teams to improve reliability, performance, and operational readiness of new featuresPartner with application engineering teams to improve service architecture, performance, and observability, and help define best practices for building reliable, scalable systemsAct as a point of support and escalation for production issuesCollaborate across teams to manage dependencies and ensure smooth system operationsTeam & Practice DevelopmentContribute to building robust SRE and operational practices across the organizationShare knowledge through documentation, pairing, and technical discussionsHelp onboard and support more junior team members as the team growsContribute to improving ways of working within the team and across EngineeringRequirementsKey qualifications and assets5-8+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform EngineeringStrong experience with cloud infrastructure (AWS, GCP, or Azure)Experience building internal platforms, tooling, or shared services that improve developer productivity and system reliabilityExperience designing systems that bridge infrastructure and application layersAbility to work across the stack: comfortable reading, debugging, and making changes to application code (e.G., backend services, APIs) when needed to improve reliability, performance, or observabilityExperience with at least one backend programming language (e.G., Node.Js, Python, Go, Java)Strong experience with monitoring, observability, and alerting tools (e.G., Datadog, Prometheus,
Grafana)Solid understanding of CI/CD systems and modern deployment practicesExperience managing infrastructure as code (e.G., Terraform, CloudFormation)Experience optimizing system performance and infrastructure costsFamiliarity with security and compliance considerations in cloud environmentsExperience working with third‑party vendors and infrastructure toolsFamiliarity with infrastructure considerations for AI/ML workloads (e.G., high‑compute services, data pipelines, or third‑party AI platforms) is a strong assetCuriosity about emerging technologies and their impact on infrastructure, reliability, and cost at scaleStrong problem‑solving skills and ability to navigate complex systemsExcellent collaboration and communication skillsNice to haveExperience contributing to platform engineering initiatives (e.G., internal developer platforms, self‑serve infrastructure)Experience improving developer experience (DX)Experience with SLIs/SLOs and reliability engineering practicesExperience mentoring or supporting other engineersOur valuesTo lead, listen first: You amplify voices that are less often heard and create space for those voices to grow. The quality of an idea doesn't correlate with the loudness of someone 's voice.The brain is a muscle: If you 're going to do something, you will do it well. Practice often and rest when needed. Give your mind what it needs to thrive.Unlearn to learn: What did we learn growing up, and what do we need to unlearn? It 's essential to understandingour personal bias and position so that we can grow.BenefitsAt Fable, you 'll join a collaborative andmission‑driven setting where you 'll work with people who caredeeply about building a more inclusive digital world. We offer benefits such as stock options, career growth opportunities, professional development support, health and dental coverage, and more.Accessibility accommodationsFable is an inclusive workplace. If you are facing any accessibility requirements or concerns regarding the hiring process or employment with us, please fill out this form or email us at
[email protected] and include the subject line "Accessibility accommodation for Senior Product Manager job application."Pay range$130,000 - $150,000 The salary band is designed to reflect the range of skills and experience needed for the position and is subject to change. The final salary is based on relevant skills, experience, and internal equity. This posting reflects an existing vacancy. Artificial intelligence (AI) tools may be used to support part of the recruitment and selection process. However, all hiring decisions are made by our hiring managers.#J-18808-Ljbffr
📌 Senior Site Reliability Engineer - $130,000 - $150,000 A Year (Toronto)
🏢 Fable
📍 Toronto