25 Sep
|
Akkodis
|
Toronto
Manager, Site Reliability Engineering (SRE)Akkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.The successful candidate will drive Site Reliability Engineering best practices, observability initiatives, automation, incident management, and continuous improvement across cloud and on-premises environments. This chance is ideal for a technical leader with strong experience in cloud operations, production support, automation, and people leadership.Work Model: HybridAbout the OpportunityAkkodis is seeking an experienced Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and scalability of enterprise applications and platforms.What Will the Successful Candidate Do?The Manager, Site Reliability Engineering (SRE)'s primary responsibilities include, but are not limited to:Lead, mentor, and develop a team of Site Reliability Engineers.Drive reliability, availability, performance, and scalability across critical applications and platforms.Implement and operationalize SRE practices, including SLIs, SLOs, error budgets, and post-incident reviews.Oversee production support operations, on-call processes, incident management, and escalations.Partner with Development, Infrastructure, Security,
and DevOps teams to improve service reliability.Drive observability and monitoring initiatives across the organization.Reduce operational effort through automation and self-healing solutions.Lead major incident response activities and support problem management processes.Support capacity planning, resiliency testing, and disaster recovery initiatives.Develop operational standards, runbooks, and knowledge management practices.Recruit, hire, onboard, and develop SRE talent.Promote a culture of continuous improvement, collaboration, and operational excellence.What the Successful Candidate Needs to SucceedMust-Have SkillsSite Reliability Engineering (SRE)Infrastructure as Code (IaC)Automation & ScriptingObservability & Monitoring ToolsProduction Support OperationsRequired Experience8+ years supporting enterprise applications and distributed systems.3+ years of experience leading technical teams.Experience implementing Site Reliability Engineering practices.Experience supporting mission-critical production environments.Experience with incident response and problem management.Experience driving automation and operational improvements.Strong stakeholder management and collaboration skills.Required QualificationsBachelor's Degree in Computer Science, Software Engineering, or equivalent experience.Strong understanding of SRE principles, SLIs, SLOs,
and error budgets.Hands-on experience with Azure and Kubernetes.Experience with Infrastructure as Code and automation frameworks.Experience with observability platforms such as Dynatrace, Datadog, Recent Relic, or AppDynamics.Strong scripting or programming skills.Excellent communication and leadership abilities.Preferred QualificationsExperience in fintech, payment processing, or regulated environments.Experience with change management and compliance practices.Knowledge of SDLC and DevOps best practices.Experience with disaster recovery and resiliency testing.Technical SkillsSite Reliability Engineering (SRE) - Must HaveInfrastructure as Code (Terraform, ARM, Bicep, etc.)Automation & ScriptingDevOps PracticesSoft SkillsStrong leadership and coaching abilitiesStrong problem-solving capabilitiesAbility to work effectively with cross-functional teamsAdditional RequirementsAbility to work in a hybrid environment in Toronto.Experience managing production support and operational teams.Ability to lead through major incidents and operational challenges.AccessibilityAt Akkodis, part of The Adecco Group, our purpose is simple: to make the future work for everyone. We foster a workplace where diversity is celebrated and every voice matters.We encourage applications from individuals of all backgrounds and identities.#SRE #SiteReliabilityEngineering #Azure #Kubernetes #CloudEngineering #EngineeringManager #DevOps #PlatformEngineering #Observability #Automation #InfrastructureAsCode #TorontoJobs #HybridJobs #TechnologyJobs #HiringNow #Akkodis #CloudOperations #ProductionSupport #LeadershipJobs
📌 Site Reliability Engineering Manager (Toronto)
🏢 Akkodis
📍 Toronto