13 Sep
|
S.i. Systems
|
Winnipeg
13 Sep
S.i. Systems
Winnipeg
Lead Site Reliability Engineering to support and enhance Azure-based developer platforms and cloud-native applications
Our financial services client is seeking a
Lead, Site Reliability Engineering (5+ years) to support and enhance Azure-based developer platforms and cloud-native applicationsJoin a team focused on improving the reliability, availability, and operability of enterprise developer platforms and internal applications. This role combines Site Reliability Engineering, platform engineering, and production support across Azure-based environments, CI/CD pipelines, observability tooling, and cloud-native services. The position supports critical application environments while contributing to automation, incident response, operational readiness, and continuous reliability improvement initiatives.Must Haves
5+ years in
Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production SupportHands-on experience with
Microsoft Azure, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Azure SQL, API Management (APIM), and Azure FunctionsProduction support experience with
incident response, troubleshooting, problem management, and operational support processesHands-on experience with
CI/CD pipelines
and deployment automation using GitHub Actions or similar platformsPost-secondary education in Computer Science, Software Engineering, Information Technology, or a related disciplineNice to Have
Experience with container apps and Kubernetes container orchestration platformsExperience supporting enterprise developer platforms or internal platform engineering teamsKnowledge of Site Reliability Engineering principles, including SLOs, SLIs, and error budgetsExposure to AI, LLM, or agent-based technology platformsAzure, Network, DevOps, or cloud-related certificationsResponsibilities
Monitor, troubleshoot,
and support applications and developer platform services across DEV, UAT, and PROD environmentsRespond to incidents and lead triage activities to restore service and resolve issuesSupport deployments, release activities, change management, and CI/CD pipelines including GitHub Actions workflowsConfigure and support Azure platform components including Azure Container Apps, Key Vault, DNS, certificates, networking, and shared cloud servicesInvestigate performance, reliability, and availability issues using Datadog, Azure Monitor, and Log AnalyticsSupport onboarding of new applications and teams to the developer platformDevelop and maintain runbooks, support procedures, knowledge articles, and operational documentationOur financial services client is seeking a
Lead, Site Reliability Engineering (5+ years) to support and enhance Azure-based developer platforms and cloud-native applicationsJoin a team focused on improving the reliability, availability, and operability of enterprise developer platforms and internal applications. This role combines Site Reliability Engineering, platform engineering, and production support across Azure-based environments, CI/CD pipelines, observability tooling, and cloud-native services. The position supports critical application environments while contributing to automation, incident response, operational readiness, and continuous reliability improvement initiatives.Permanent, Toronto, 4 days/ week on siteMust Haves
5+ years in
Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production SupportHands-on experience with
Microsoft Azure, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Azure SQL, API Management (APIM), and Azure FunctionsProduction support experience with
incident response, troubleshooting, problem management, and operational support processesHands-on experience with
CI/CD pipelines
and deployment automation using GitHub Actions or similar platformsPost-secondary education in Computer Science, Software Engineering, Information Technology, or a related disciplineNice to Have
Experience with container apps and Kubernetes container orchestration platformsExperience supporting enterprise developer platforms or internal platform engineering teamsKnowledge of Site Reliability Engineering principles, including SLOs, SLIs, and error budgetsExposure to AI, LLM, or agent-based technology platformsAzure, Network, DevOps, or cloud-related certificationsResponsibilities
Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environmentsRespond to incidents and lead triage activities to restore service and resolve issuesSupport deployments, release activities, change management, and CI/CD pipelines including GitHub Actions workflowsConfigure and support Azure platform components including Azure Container Apps, Key Vault, DNS, certificates, networking, and shared cloud servicesInvestigate performance, reliability, and availability issues using Datadog, Azure Monitor, and Log AnalyticsSupport onboarding of recent applications and teams to the developer platformDevelop and maintain runbooks, support procedures, knowledge articles, and operational documentation
#J-18808-Ljbffr
📌 Lead Site Reliability Engineering To Support And Enhance Azure-Based Developer Platforms And Cloud-N (Winnipeg)
🏢 S.i. Systems
📍 Winnipeg