- Operate and maintain the Kubernetes based platform across public and private cloud environments.
- Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.
- Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews.
- Onboard and support new clients onto the platform.
- Build automation and diagnostic tooling that cuts manual effort and evaluates performance.
- Identify and deliver process improvements.
- Support software and hardware upgrades and keep components up to date.
- Proactively manage capacity so the platform scales with demand.
Required Skills:
- Hands on experience operating Kubernetes in production,
ideally with Service Mesh.
- Robust Linux and command line fundamentals.
- Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure.
- Experience working with a public cloud provider, preferable Azure or AWS.
- Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus.
- Scripting or coding in Python or Java is a strong plus.
- CI/CD, infrastructure as code such as Helm or Terraform is a plus
- A financial services background is not required.
📌 Site Reliability Engineer (Montreal)
🏢 Open Systems Technologies
📍 Montreal
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.