23 Aug
|
Open Systems Technologies
|
Montreal
23 Aug
Open Systems Technologies
Montreal
- Operate and maintain the Kubernetes based platform across public and private cloud environments.
- Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.
- Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews.
- Onboard and support new clients onto the platform.
- Build automation and diagnostic tooling that cuts manual effort and evaluates performance.
- Identify and deliver process improvements.
- Support software and hardware upgrades and keep components up to date.
- Proactively manage capacity so the platform scales with demand.
Required Skills:
- Hands on experience operating Kubernetes in production, ideally with Service Mesh.
- Strong Linux and command line fundamentals.
- Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure.
- Experience working with a public cloud provider, preferable Azure or AWS.
- Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus.
- Scripting or coding in Python or Java is a robust plus.
- CI/CD, infrastructure as code such as Helm or Terraform is a plus
- A financial services background is not required.
📌 Site Reliability Engineer (Montreal)
🏢 Open Systems Technologies
📍 Montreal