Drive proactive reliability at IBM Software, focusing on incident command in multi-cloud settings. Leverage hands-on engineering skills to enhance automation and improve incident management practices.IBM Software seeks a seasoned Reliability Engineer with extensive experience in incident management and SRE. This role involves 75% technical work, including designing reliability improvements, and 25% coaching on incident response. You'll work closely with engineering leaders to elevate reliability across all teams.Key Responsibilities:
Analyze failure patterns and propose reliability enhancements
Manage Rootly configurations and integration with key tools
Define SLO/SLA frameworks and guide reliability investments
Foster continuous improvement of incident response processes
Develop and deliver training programs on incident managementRequirements:
10+ years in SRE or reliability engineering
Expertise in AWS, GCP, or Azure
Proficient with incident management tools like Rootly
Solid written communication skills
Experience in large organizations with 500+ engineersElevate your career by improving incident management and driving reliability initiatives at IBM Software.#J-18808-Ljbffr
📌 Reliability Engineer At Ibm Cloud Toronto
🏢 IBM
📍 Toronto