11 Sep
|
DataPattern
|
Montreal
11 Sep
DataPattern
Montreal
Platform & Site Reliability Engineer
Location: Montreal, QC
Work Model: Hybrid – 3 days onsite per week
Experience: 6+ Years
Employment Type: Contract
Position Overview
We are seeking an experienced Platform & Site Reliability Engineer (SRE) with 6+ years of experience in platform engineering, cloud-native technologies, container orchestration, infrastructure automation, and site reliability practices.
The ideal candidate will have strong hands-on experience with Kubernetes, Docker, CI/CD, observability, Linux, and automation. You will work closely with application and infrastructure teams to build reliable platforms, automate deployments, improve system availability, and troubleshoot production issues.
Key Responsibilities
- Design, deploy, and manage scalable, containerized applications using Kubernetes and Docker.
- Build, maintain, and optimize CI/CD pipelines for reliable and repeatable deployments.
- Implement and maintain platform observability, monitoring, metrics, and alerting using tools such as Prometheus and Grafana.
- Collaborate with application teams to onboard services and implement authentication and authorization mechanisms such as OAuth2, SSO, and JWT.
- Troubleshoot production issues, perform root cause analysis (RCA), and drive improvements in system reliability and availability.
- Automate routine operational tasks using Python, Ansible, or similar scripting/configuration tools.
- Support web-based application deployments, including APIs, UI applications, and application configurations.
- Participate in system design, capacity planning, performance tuning, and platform optimization.
- Establish and maintain best practices for deployment, monitoring, reliability, and operational readiness.
- Collaborate with development,
infrastructure, and other cross-functional teams to continuously improve platform stability.
Must-Have Skills
- 6+ years of experience in Platform Engineering, SRE, DevOps, or related roles.
- Solid hands-on experience with Kubernetes and Docker.
- Strong understanding of containerization, orchestration, and deployment strategies.
- Hands-on experience with observability and monitoring tools such as:
- Prometheus
- Grafana
- Metrics and Alerting
- Good understanding of web deployment and web application/component architecture.
- Experience with authentication and authorization, including:
- OAuth/OAuth2
- SSO
- JWT
- Strong understanding of CI/CD pipelines and deployment automation.
- Experience with tools such as Jenkins, GitLab CI, or similar.
- Strong Linux administration and troubleshooting skills.
- Basic knowledge of SQL and database operations.
Nice-to-Have Skills
- Strong Python scripting experience for automation.
- Experience with Ansible or other configuration management tools.
- Experience with Infrastructure as Code (IaC) tools such as Terraform.
- Experience with Helm and Kubernetes package management.
- Exposure to cloud platforms and cloud-native services.
- Experience with infrastructure automation and platform standardization.
Preferred Qualifications
- Strong troubleshooting and problem-solving skills.
- Understanding of SRE principles, reliability engineering, and production support.
- Experience working in fast-paced, cross-functional environments.
- Strong communication and collaboration skills.
- Ability to work independently while partnering effectively with development and infrastructure teams.
- Passion for automation, system stability, observability, and continuous improvement.
📌 Platform & Site Reliability Engineer (Montreal)
🏢 DataPattern
📍 Montreal