Join Johnson Controls as a Staff Site Reliability Engineer focusing on the OpenBlue platform in Canada. Lead the charge in ensuring system reliability and addressing critical production challenges.
This senior engineering role is designed for seasoned professionals who will manage escalated issues across our data platforms and provide technical leadership during major incidents. You will be responsible for implementing infrastructure as code with Terraform and enhancing cloud operations across Azure and AWS. Your mission is to foster operational stability through root cause analysis and monitoring improvements.
Key Responsibilities:
• Oversee escalated issues for OpenBlue Data Platform
• Conduct comprehensive debugging of production failures
• Drive permanent fixes through effective analysis
• Improve system monitoring and alerting strategies
• Lead incident resolution as a senior technical contact
Requirements:
• Minimum of 7 years in site reliability engineering
• Expertise in Terraform and cloud operations
• Practical experience with Kubernetes at scale
• Proficient in using Datadog for monitoring
• Able to respond to incidents outside business hours
Your expertise will directly impact our customer service and operational integrity at Johnson Controls.
#J-18808-Ljbffr