Join IBM Software as a Senior Incident Command Developer, enhancing reliability in a cloud-native setting. Focus on engineering day-to-day while teaching teams incident management best practices.
In this role, you will contribute to critical incident management and reliability tasks within IBM’s Cloud Architecture and Reliability team. You will spend 75% of your time on hands-on engineering duties, like automating processes and improving tooling, while dedicating 25% to team coaching and post-mortem analysis.
Key Responsibilities:
• Identify and mitigate systemic failure patterns
• Own incident management tooling and workflows
• Define SLO/SLA frameworks and monitor performance
• Review incident documentation for clarity and quality
• Conduct training sessions to enhance team capabilities
Requirements:
• 10+ years in SRE, incident management, or reliability
• Proficient in AWS, GCP, or Azure environments
• Deep experience with reliability tooling
• Strong understanding of distributed systems
• Effective communication and coaching skills
Become a leader in incident management and reliability at IBM Software, shaping the future of cloud solutions.
#J-18808-Ljbffr
📌 Senior Incident Command Developer IBM (Ontario)
🏢 IBM
📍 Ontario