Drive proactive reliability at IBM Software, focusing on incident command in multi-cloud environments. Leverage hands-on engineering skills to enhance automation and improve incident management practices.
IBM Software seeks a seasoned Reliability Engineer with extensive experience in incident management and SRE. This role involves 75% technical work, including designing reliability improvements, and 25% coaching on incident response. You'll work closely with engineering leaders to elevate reliability across all teams.
Key Responsibilities:
• Analyze failure patterns and propose reliability enhancements
• Manage Rootly configurations and integration with key tools
• Define SLO/SLA frameworks and guide reliability investments
• Foster continuous improvement of incident response processes
• Develop and deliver training programs on incident management
Requirements:
• 10+ years in SRE or reliability engineering
• Expertise in AWS, GCP, or Azure
• Proficient with incident management tools like Rootly
• Robust written communication skills
• Experience in large organizations with 500+ engineers
Elevate your career by improving incident management and driving reliability initiatives at IBM Software.
#J-18808-Ljbffr
📌 Reliability Engineer at IBM Cloud (Ontario)
🏢 IBM
📍 Ontario
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.