Site Reliability Engineer Kubernetes Ontario (Canada)

Site Reliability Engineer Kubernetes Ontario (Canada)

31 Jul
|
Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision!
|
Canada

31 Jul

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision!

Canada

Enterprise Kubernetes SRE (Python, GitOps, API, Container, Cloud,, MongoDB, Postgres)
Toronto, ON - Hybrid (4 Days WFO)
12 months
We are seeking an experienced Site Reliability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll be responsible for ensuring the reliability, performance, and scalability of our enterprise-grade Kubernetes platform that powers mission-critical applications across the organization.
This role places a robust emphasis on automation, with the expectation that the successful candidate will continuously identify and eliminate manual toil through intelligent tooling, self-healing systems, and AI-assisted operational workflows. You will work alongside platform engineers, DevOps teams, and application developers to build and maintain a world-class container orchestration platform.
WHAT YOU'LL DO PLATFORM RELIABILITY & OPERATIONS Ensure 99.9% availability SLA for platform services across 60+ Kubernetes clusters spanning production, DR, UAT, QA, and development environments
Manage and operate enterprise Kubernetes distributions across on-premises and cloud-hosted environments




Implement and maintain disaster recovery patterns across multi-AZ architectures and geographically distributed data centres
Design and execute capacity planning, resource optimization, and cluster scaling strategies
Support full cluster lifecycle operations including provisioning, upgrades, patching, and decommissioning
Automate cluster health checks and validation workflows for continuous reliability assurance
Manage multi-tenant cluster settings with strict isolation and RBAC enforcement
INCIDENT MANAGEMENT & ON-CALL Participate in on-call rotation for platform infrastructure support with sub-15-minute MTTR targets
Lead incident response, troubleshooting, and root cause analysis for platform issues
Conduct blameless post-incident reviews and implement preventive measures
Develop and maintain runbooks, troubleshooting guides, and operational playbooks
Coordinate with application teams during incidents affecting workloads
Build automated incident detection and response systems to reduce manual intervention
Integrate AI-assisted triage tools for faster incident classification and resolution
J-18808-Ljbffr

📌 Site Reliability Engineer Kubernetes Ontario (Canada)
🏢 Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision!
📍 Canada

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer kubernetes ontario (canada) / canada

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer kubernetes ontario (canada) / canada