Responsibilities
Operate and maintain the Kubernetes based platform across public and private cloud environments.
Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.
Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews.
Onboard and support current clients onto the platform.
Build automation and diagnostic tooling that cuts manual effort and evaluates performance.
Identify and deliver process improvements.
Support software and hardware upgrades and keep components up to date.
Proactively manage capacity so the platform scales with demand.
Required Skills:
Hands on experience operating Kubernetes in production,
ideally with Service Mesh.
Robust Linux and command line fundamentals.
Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure.
Experience working with a public cloud provider, preferable Azure or AWS.
Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus.
Scripting or coding in Python or Java is a robust plus.
CI/CD, infrastructure as code such as Helm or Terraform is a plus
A financial services background is not required.
📌 Site Reliability Engineer Montreal (Canada)
🏢 Open Systems Technologies
📍 Canada
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.