Site Reliability Engineer (Montreal)

Site Reliability Engineer (Montreal)

15 Aug
|
Mantu
|
Montreal

15 Aug

Mantu

Montreal

About the Role

Join an API Platform team responsible for operating and continuously improving a critical enterprise platform used by development teams across the organization. The team follows an API-first approach and supports a diverse technology landscape spanning public cloud, hybrid cloud, and on-premises environments.

This is a hands-on SRE / Platform Engineering opportunity that combines software engineering and operations. You will help keep the Kubernetes platform stable, secure, performant, and scalable while building automation, improving observability, supporting production incidents, and enhancing the overall developer experience. The role is focused on the day-to-day operation and continuous improvement of the platform, rather than traditional application development.

Main Responsibilities

- Operate and maintain a Kubernetes-based platform across public cloud, private cloud, and on-premises environments.
- Monitor platform health using observability tools, ensuring alerts are actionable and supported by accurate runbooks and documentation.
- Manage incidents end-to-end, including incident response, troubleshooting, root cause analysis, and post-incident reviews.
- Build automation,



diagnostic tools, and performance tests to reduce manual effort and improve platform reliability.
- Support client onboarding, platform upgrades, infrastructure changes, and capacity management to ensure the platform scales with demand.
- Identify and implement continuous improvements across operations, automation, deployment, monitoring, and support processes.

Qualifications

- 5–7 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or a similar infrastructure-focused role.
- Solid hands-on experience operating Kubernetes in production, ideally with Service Mesh exposure.
- Excellent Linux and command-line skills with the ability to troubleshoot complex systems from application to infrastructure layers.
- Experience with a major public cloud platform, preferably Azure or AWS.
- Familiarity with Grafana, Prometheus, Loki, and Tempo; Python or Java scripting experience is an asset.
- Knowledge of CI/CD and Infrastructure as Code, particularly Helm or Terraform, combined with strong analytical and problem-solving skills.

📌 Site Reliability Engineer (Montreal)
🏢 Mantu
📍 Montreal

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (montreal) / montreal

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (montreal) / montreal