Make a significant impact at Confluent as an Expert Site Reliability Engineer focused on incident management and reliability enhancements. You'll work within a multi-cloud architecture to optimize performance and reliability.This expert role blends 75% technical engineering with 25% strategy, involving the analysis of systemic failure patterns, designing reliability frameworks, and teaching best practices. You'll be instrumental in developing incident response processes that facilitate organizational success and sustainability. Join a global team dedicated to improving cloud-based reliability.Key Responsibilities:
- Analyze and improve systemic failure patterns
- Own configuration and workflows for incident management tools
- Define SLO/SLA frameworks to guide reliability investments
- Edit incident documents for customer clarity
- Lead training programs and coach teams through post-mortemsRequirements:
- 10+ years of experience in SRE or incident management
- Cloud experience with AWS, GCP, or Azure
- Expertise in incident management tools such as Rootly
- Solid understanding of distributed systems
- Experience in cultural change within engineering organizationsUtilize your reliability engineering expertise to drive impactful changes across Confluent's architecture.#J-18808-Ljbffr
📌 Expert Site Reliability Engineer At Confluent (Toronto)
🏢 IBM
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.