What is the Opportunity? Client is seeking to hire a Senior Site Reliability Engineer for its Application Maintenance and Transformation| Data Services and Integration team. As a Senior Site Reliability Engineer| you will bring the engineering mindset of bold ambition| curiosity and outcome focus to ensuring the performance and reliability of our systems.
This role calls for a energetic individual who excels in a collaborative environment| interacting with cross-functional teams to establish best practices for observability| monitoring| logging| alerting| and automation. What will you do?Set vision for SRE product base (monitoring| alerting| self-healing| reliability testing).Lead cross-functional collaborations to define and implement best practices for monitoring| logging| and incident response| driving a proactive stance on system health.Function as portfolio SME (Subject Matter Expert)understand & document common components| core functionalities| infrastructure of supported applications.Leverage AI-assisted operational tools and platforms to improve incident detection| root cause analysis| alert correlation| capacity forecasting| and overall platform reliability.Implement and support AIOps practices by utilizing AI-driven observability and monitoring solutions to proactively identify anomalies| reduce MTTR| and enhance platform resilience.Drive adoption of Generative AI and AI-powered engineering tools to automate operational workflows| runbook generation| knowledge management| troubleshooting| and production support activities.Actively participate in deploying software applications| automation tools| and IT infrastructure.Work closely with development teams to understand code changes and their impact on the production environment| ensuring that new releases meet our reliability standards.Drive transformation by continuously looking for ways to automate existing SRE processes and increase operational efficiency.Guide the technical direction for future deployments| advocating for reliability and performance improvements based on industry trends and company objectives.Lead in incident management and problem management for applications in scope and RCA action items fulfillment/ownership.Debug production issues across services and levels of the stack and provide primary operational support.Perform occasional off-hours support.
📌 Release Lead (Ontario)
🏢 Quantum World Technologies
📍 Ontario
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.