Make a significant impact at Confluent as an Expert Site Reliability Engineer focused on incident management and reliability enhancements. You'll work within a multi-cloud architecture to optimize performance and reliability.
This expert role blends 75% technical engineering with 25% strategy, involving the analysis of systemic failure patterns, designing reliability frameworks, and teaching best practices. You'll be instrumental in developing incident response processes that facilitate organizational success and sustainability. Join a global team dedicated to improving cloud-based reliability.
Key Responsibilities:
• Analyze and improve systemic failure patterns
• Own configuration and workflows for incident management tools
• Define SLO/SLA frameworks to guide reliability investments
• Edit incident documents for customer clarity
• Lead training programs and coach teams through post-mortems
Requirements:
• 10+ years of experience in SRE or incident management
• Cloud experience with AWS, GCP, or Azure
• Expertise in incident management tools such as Rootly
• Robust understanding of distributed systems
• Experience in cultural change within engineering organizations
Utilize your reliability engineering expertise to drive impactful changes across Confluent's architecture.
#J-18808-Ljbffr
📌 Expert Site Reliability Engineer at Confluent (Toronto)
🏢 IBM
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.