08 Sep
|
Recutify
|
Toronto
Position: Incident Manager
Location: Toronto, Canada
Mode of Work: Hybrid ( 3 or 4 days office in downtown)
Type: FTE
This job posting is for an existing, active vacancy and we are looking to hire an Incident Manager
Summary :
We are looking for an experienced Incident Manager responsible for end‑to‑end ownership of production incidents across business‑critical platforms. The role demands strong production support expertise, advanced SQL‑based investigation capability, and the ability to leverage Generative AI tools for faster triage, RCA acceleration, and operational efficiency.
Roles & Responsibilities :
- Incident Management & Operational Leadership
- Production Support & Deep Technical Analysis:
Perform hands‑on production investigation using logs, monitoring tools, batch/job controls, and application telemetry. Use advanced SQL for:
– Root cause investigation
– Data validation
– Large‑volume analysis and performance tuning
- GenAI‑Enabled Incident Operations including AI‑assisted triage, automated RCA generation, and insights extraction from historical ticket and knowledge repositories
- Ensure rapid service restoration, accurate impact assessment, and prevention of incident recurrence
- Collaborate with distributed teams to drive stability, continuous improvement, and operational excellence
Qualifications & Skills :
Core Experience
- 8–12+ years in Production Support / Incident Management within large‑scale enterprise environments
- Proven experience managing P1/P2 incidents in mission‑critical systems
- Robust understanding of ITIL / ITSM processes, preferably within banking or regulated environments
SQL & Technical Expertise (Must Have)
- Advanced SQL proficiency (joins, subqueries, performance tuning, large‑data analysis)
- Strong experience in production data investigation and validation
- Ability to convert data insights into clear technical & business communication
GenAI & Automation Exposure
- Hands‑on experience or strong exposure to Generative AI in IT Operations, including:
- 1. AI‑assisted log analysis
2. AI‑generated RCA and incident summaries
3. Knowledge mining from historical tickets & documentation
- Familiarity with AI‑enabled ITSM tools (e.g., ServiceNow Now Assist or similar)
Tools & Platforms
- ServiceNow (Incident, Problem, Change modules)
- Monitoring & logging tools: ELK, Splunk, AppDynamics, Dynatrace (or equivalent)
- Batch processing, data platforms, and distributed system architectures
Nice to Have :
- Experience in Banking, Wealth Management, Insurance, or Capital Markets
- Experience supporting data‑heavy platforms (batch, ETL, analytics, reporting)
- Exposure to AIOps / predictive operations
- Experience working with hybrid onshore/offshore support models
📌 Incident Manager (Toronto)
🏢 Recutify
📍 Toronto