25 Aug
|
Quantum Technology Recruiting Inc. (QTR)
|
Winnipeg
25 Aug
Quantum Technology Recruiting Inc. (QTR)
Winnipeg
Position:
Manager, Site Reliability Engineering (SRE) Location:
Toronto Salary:
$155,000 - $165,000 Posting Type:
Open vacancy Our client is looking for an experienced
Manager, Site Reliability Engineering (SRE)
to lead a team responsible for the reliability, availability, performance, and resiliency of business-critical applications and platforms. This is a newly created position reporting to the Director, SRE & DevOps. The successful candidate will combine people leadership with strong hands-on technical expertise, helping mature an evolving SRE function while supporting a large portfolio of critical applications. Our client is looking for someone who can lead and develop an SRE team while remaining technically credible and hands-on across cloud, automation, observability, infrastructure and production reliability. Description
Lead, coach and develop a team of Site Reliability Engineers. Own the reliability, availability, scalability and performance of assigned business-critical applications and platforms. Help mature SRE practices including SLIs, SLOs, error budgets, incident response, post-incident reviews and continuous reliability improvement. Oversee production operations, incident management, escalation handling and problem management. Drive improvements in observability, monitoring and operational visibility. Partner with Development, DevOps, Infrastructure, Security and Incident Management teams. Reduce operational toil through automation and self-healing capabilities. Lead capacity planning,
resiliency testing and disaster recovery readiness. Act as a senior technical escalation point during major production incidents. Establish effective runbooks, documentation standards and knowledge-management practices. Help develop reliability roadmaps for the applications within your portfolio. Support hiring and workforce planning as the SRE organization continues to evolve. Skills
8+ years of experience in senior technical roles supporting complex or distributed systems. 3+ years of people leadership experience. Strong hands-on knowledge of at least one major cloud platform, with Azure preferred; AWS is also acceptable. Strong knowledge of Kubernetes. Near-expert capability with Infrastructure as Code, automation and observability. Strong experience with enterprise observability tools; Dynatrace is preferred, with Datadog, New Relic, AppDynamics or similar also relevant. Linux knowledge – this is an key technical requirement. Strong scripting and/or programming capabilities. Strong understanding of SRE principles including SLOs, SLIs, incident management, resiliency and toil reduction. Ability to operate as both a people leader and hands-on technical leader. Excellent communication and stakeholder-management skills. Nice to Have
Experience within payments, fintech or other highly regulated environments. Exposure to PCI DSS and/or NIST frameworks. Experience working with change-management and compliance processes. Familiarity with SDLC best practices.
#J-18808-Ljbffr
📌 Manager, Site Reliability Engineering (SRE) (Winnipeg)
🏢 Quantum Technology Recruiting Inc. (QTR)
📍 Winnipeg