Staff Site Reliability Engineer - Confluent Incident Management & Reliability (Toronto)

Staff Site Reliability Engineer - Confluent Incident Management & Reliability (Toronto)

09 Aug
|
IBM
|
Toronto

09 Aug

IBM

Toronto

Your Role and ResponsibilitiesAbout the RoleConfluent Cloud processes millions of events per second across AWS, GCP, and Azure. When incidents happen in a multi‐cloud streaming platform, they happen at scale—data in motion, exactly‐once semantics, and cascading failure modes that require deep systems thinking. We need an expert‐level engineer who can drive proactive reliability improvements that prevent these incidents before they occur.This role combines hands‐on technical work with strategic program ownership. You'll spend roughly 75% of your time on engineering: building automation, improving tooling, analyzing systemic failure patterns, and designing reliability improvements. The remaining 25% is teaching and coordination: coaching teams through post‐mortems, training incident commanders, and evolving our incident response practices.You'll be part of a global team with follow‐the‐sun coverage, with clean handoffs that keep everyone working sustainable hours. This role sits within Cloud Architecture and Reliability - Supportability, a horizontal team that owns reliability standards and tooling across engineering. You're the person who makes us need incident management less.What You Will DoAnalyze systemic failure patterns and design reliability improvements that prevent incident recurrenceOwn Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and SlackDefine and maintain SLO/SLA frameworks; use error budgets to guide reliability investmentsOwn standards, practices,



and continuous improvement of incident response across engineeringEdit and review customer‐facing incident documents (CRCAs) to ensure quality and clarityDevelop and deliver training programs; coach teams through post‐mortemsPartner with engineering leaders to elevate reliability practices org‐wideDeep experience with observability: metrics, logging, tracingKubernetes and container orchestration experienceUnderstanding of CI/CD pipelines and release processesStrong written communication (design docs, runbooks, post‐mortems)Experience driving org‐wide process and cultural changesPreferred EducationMaster's DegreeRequired Technical And Professional Expertise10+ years of relevant experience in SRE, incident management, or reliability engineeringCloud experience with at least one of AWS, GCP, or Azure (we run all three)Experience navigating reliability/incident programs at 500+ engineer organizationsDeep expertise with incident management tooling (Rootly, PagerDuty, or similar)Strong understanding of distributed systems and failure modes at scaleKafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systemsPreferred Technical And Qualified ExperienceAdvanced Cloud Knowledge: Experience with cloud-based infrastructure and its application in reliability and resiliency engineering.Specialized Scripting Skills: Proficiency in scripting languages and automation tools to optimize system reliability and performance. #J-18808-Ljbffr

📌 Staff Site Reliability Engineer - Confluent Incident Management & Reliability (Toronto)
🏢 IBM
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: staff site reliability engineer - confluent incident management & reliability (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: staff site reliability engineer - confluent incident management & reliability (toronto) / toronto