Site Reliability Engineer - Kubernetes (Toronto)

Site Reliability Engineer - Kubernetes (Toronto)

31 Jul
|
Astra North Infoteck
|
Toronto

31 Jul

Astra North Infoteck

Toronto

Job Description Enterprise Kubernetes SRE (Python, Git

Ops, API, Container, Cloud,, MongoDB, Postgres) Toronto, ON - Hybrid (4 Days WFO) 12 months We are seeking an experienced Site Reliability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization.

You''ll be responsible for ensuring the reliability, performance, and scalability of our enterprise-grade Kubernetes platform that powers mission-critical applications across the organization.

This role places a robust emphasis on automation, with the expectation that the successful candidate will continuously identify and eliminate manual toil through intelligent tooling, self-healing systems, and AI-assisted operational workflows.

You will work alongside platform engineers, Dev

Ops teams, and application developers to build and maintain a world-class container orchestration platform. ====================================================================== WHAT YOU''LL DO ====================================================================== PLATFORM RELIABILITY & OPERATIONS -------------------------------------------------------------------------------- Ensure 99.9% availability SLA for platform services across 60+ Kubernetes clusters spanning production, DR, UAT, QA,



and development environments Manage and operate enterprise Kubernetes distributions across on-premises and cloud-hosted environments Implement and maintain disaster recovery patterns across multi-AZ architectures and geographically distributed data centres Design and execute capacity planning, resource optimization, and cluster scaling strategies Support full cluster lifecycle operations including provisioning, upgrades, patching, and decommissioning Automate cluster health checks and validation workflows for continuous reliability assurance Manage multi-tenant cluster environments with strict isolation and RBAC enforcement INCIDENT MANAGEMENT & ON-CALL -------------------------------------------------------------------------------- Participate in on-call rotation for platform infrastructure support with sub-15-minute MTTR targets Lead incident response, troubleshooting, and root cause analysis for platform issues Conduct blameless post-incident reviews and implement preventive measures Develop and maintain runbooks, troubleshooting guides, and operational playbooks Coordinate with application teams during incidents affecting workloads Build automated incident detection and response systems to reduce manual intervention Integrate AI-assisted triage tools for faster incident classification and resolution

📌 Site Reliability Engineer - Kubernetes (Toronto)
🏢 Astra North Infoteck
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer - kubernetes (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer - kubernetes (toronto) / toronto