Make a significant impact at Confluent as an Expert Site Reliability Engineer focused on incident management and reliability enhancements. You'll work within a multi-cloud architecture to optimize performance and reliability.This expert role blends 75% technical engineering with 25% strategy, involving the analysis of systemic failure patterns, designing reliability frameworks, and teaching best practices. You'll be instrumental in developing incident response processes that facilitate organizational success and sustainability.
Join a global team dedicated to improving cloud-based reliability.Key Responsibilities:Analyze and improve systemic failure patternsOwn configuration and workflows for incident management toolsDefine SLO/SLA frameworks to guide reliability investmentsEdit incident documents for customer clarityLead training programs and coach teams through post-mortemsRequirements:10+ years of experience in SRE or incident managementCloud experience with AWS, GCP, or AzureExpertise in incident management tools such as RootlyStrong understanding of distributed systemsExperience in cultural change within engineering organizationsUtilize your reliability engineering expertise to drive impactful changes across Confluent's architecture.
📌 Expert Site Reliability Engineer At Confluent (Toronto)
🏢 IBM
📍 Toronto