We are seeking a dynamic individual to join our Application Maintenance and Transformation | Data Services and Integration team. This role requires a cooperative approach, working with cross-functional teams to establish and implement best practices for system reliability and performance.
Required Skills & Qualifications
- Expertise in Site Reliability Engineering (SRE), DevOps, Kubernetes/OpenShift, and cloud-native applications
- Strong skills in observability, monitoring, logging, alerting, and automation
- Experience with AIOps and AI-powered tools, including AI/ML workloads and LLMs
- Proficiency in incident management and problem-solving
- Prior work experience in the client's industry
Preferred Skills & Qualifications
- Experience with AI-assisted operational tools and platforms
- Knowledge of Generative AI and AI-powered engineering tools
- Background in automating operational workflows and production support activities
Day-to-Day Responsibilities
- Set vision for SRE product base, including monitoring, alerting, self-healing, and reliability testing
- Lead cross-functional collaborations to implement best practices for monitoring, logging, and incident response
- Function as a portfolio Subject Matter Expert (SME) to understand and document core functionalities and infrastructure
- Drive transformation by automating SRE processes and increasing operational efficiency
- Guide technical direction for future deployments, advocating for reliability and performance improvements
- Lead in incident and problem management, including RCA action items fulfillment
- Debug production issues and provide primary operational support
- Perform occasional off-hours support
📌 Senior Site Reliability Engineer (Toronto)
🏢 Artech
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.