Site Reliability Engineering Manager (Toronto)

Site Reliability Engineering Manager (Toronto)

18 Sep
|
Akkodis
|
Toronto

18 Sep

Akkodis

Toronto

Manager, Site Reliability Engineering (SRE)

Toronto, Ontario (Hybrid)

Permanent Opportunity

Location: Toronto, Ontario

Work Model: Hybrid

About the Opportunity

Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.

The successful candidate will drive Site Reliability Engineering best practices, observability initiatives, automation, incident management, and continuous improvement across cloud and on-premises environments. This opportunity is ideal for a technical leader with strong experience in cloud operations, production support, automation, and people leadership.

What Will the Successful Candidate Do?

The Manager, Site Reliability Engineering (SRE)'s primary responsibilities include, but are not limited to:

- Lead, mentor, and develop a team of Site Reliability Engineers.
- Drive reliability, availability, performance, and scalability across critical applications and platforms.
- Implement and operationalize SRE practices, including SLIs, SLOs, error budgets, and post-incident reviews.
- Oversee production support operations, on-call processes, incident management, and escalations.
- Partner with Development, Infrastructure, Security, and DevOps teams to improve service reliability.
- Drive observability and monitoring initiatives across the organization.
- Reduce operational effort through automation and self-healing solutions.
- Lead major incident response activities and support problem management processes.
- Support capacity planning, resiliency testing, and disaster recovery initiatives.
- Develop operational standards, runbooks, and knowledge management practices.
- Recruit, hire, onboard, and develop SRE talent.




- Promote a culture of continuous improvement, collaboration, and operational excellence.

What the Successful Candidate Needs to Succeed

Must-Have Skills

- Site Reliability Engineering (SRE)
- Azure Cloud
- Kubernetes
- Infrastructure as Code (IaC)
- Linux
- Automation & Scripting
- Observability & Monitoring Tools
- Incident Management
- Production Support Operations

Required Experience

- 8+ years supporting enterprise applications and distributed systems.
- 3+ years of experience leading technical teams.
- Experience implementing Site Reliability Engineering practices.
- Experience supporting mission-critical production environments.
- Experience with incident response and problem management.
- Experience driving automation and operational improvements.
- Strong stakeholder management and collaboration skills.

Required Qualifications

- Bachelor’s Degree in Computer Science, Software Engineering, or equivalent experience.
- Robust understanding of SRE principles, SLIs, SLOs, and error budgets.
- Hands-on experience with Azure and Kubernetes.
- Experience with Infrastructure as Code and automation frameworks.
- Experience with observability platforms such as Dynatrace, Datadog, New Relic, or AppDynamics.
- Strong scripting or programming skills.
- Excellent communication and leadership abilities.

Preferred Qualifications

- Experience in fintech, payment processing, or regulated environments.




- Experience with change management and compliance practices.
- Knowledge of SDLC and DevOps best practices.
- Experience with disaster recovery and resiliency testing.

Technical Skills

- Site Reliability Engineering (SRE) - Must Have
- Azure Cloud Platform
- Kubernetes
- Infrastructure as Code (Terraform, ARM, Bicep, etc.)
- Linux
- Dynatrace, Datadog, New Relic, or AppDynamics
- Automation & Scripting
- Incident Management
- DevOps Practices

Soft Skills

- Strong leadership and coaching abilities
- Excellent communication skills
- Strong problem-solving capabilities
- Ability to work effectively with cross-functional teams
- Continuous improvement mindset

Additional Requirements

- Ability to work in a hybrid environment in Toronto.
- Experience managing production support and operational teams.
- Strong stakeholder-facing experience.
- Ability to lead through major incidents and operational challenges.

How to Apply

Submit your resume in confidence or apply through the Akkodis Canada website.

We thank all applicants for their interest in this opportunity. Only candidates who meet the qualifications outlined above will be contacted for further discussions.

Accessibility

At Akkodis, part of The Adecco Group, our purpose is simple: to make the future work for everyone. We foster a workplace where diversity is celebrated and every voice matters.

We encourage applications from individuals of all backgrounds and identities.

#SRE #SiteReliabilityEngineering #Azure #Kubernetes #CloudEngineering #EngineeringManager #DevOps #PlatformEngineering #Observability #Automation #InfrastructureAsCode #TorontoJobs #HybridJobs #TechnologyJobs #HiringNow #Akkodis #CloudOperations #ProductionSupport #LeadershipJobs

📌 Site Reliability Engineering Manager (Toronto)
🏢 Akkodis
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineering manager (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineering manager (toronto) / toronto