22 Sep
|
AceStack
|
Montreal
Job Title
Core SRE Engineer Job Type:
Full Time Location
Montreal, QC (Hybrid)
We are seeking an experienced
Core SRE Engineer with 8+ years of IT experience and strong expertise in
Linux/Unix, Python, Shell Scripting, monitoring, application support, and site reliability engineering . The ideal candidate will have hands-on experience supporting enterprise middleware and BI platforms, managing production environments, troubleshooting complex incidents, and improving system reliability through automation.
The role involves
L3/global production support , infrastructure management, incident and problem management, automation, and close collaboration with engineering and development teams.
Key Responsibilities
Provide Level 3 SRE and production support for enterprise middleware and application platforms.
Manage and support applications and tooling involving
Apache ZooKeeper, Ansible Automation Platform, and Terraform .
Act as the highest level of escalation for production incidents and service stability issues.
Troubleshoot complex Linux/Unix, application, middleware, VMware, and load-balancer issues.
Monitor application and infrastructure health using
Splunk, Grafana, Prometheus, and Loki .
Participate in incident, change, escalation, and problem management activities.
Collaborate with engineering and development teams to resolve production issues and improve service reliability.
Automate operational processes using
Python, Shell scripting, Ansible, and Terraform to reduce manual effort and operational toil.
Support code releases and coordinate with development teams during application deployments.
Manage production escalations and participate in on-call/weekend support as required.
Prepare and submit operational reports and coordinate with multiple stakeholders.
Maintain strong documentation and ensure effective knowledge transfer across global teams.
Required Skills & Qualifications
8+ years of overall IT experience, with 8+ years in SRE or a similar production support role.
Advanced hands-on experience with
Linux/Unix administration and support.
Strong
Shell scripting and Python programming skills.
Experience with
Splunk and/or Grafana, Prometheus, and Loki monitoring stacks.
Working knowledge of
Veritas Cluster Service, Load Balancers, and VMware .
Strong understanding of
ITIL principles and IT service management practices.
Experience supporting BI platforms and enterprise middleware environments.
Strong troubleshooting and outage-management capabilities.
Excellent written and verbal communication skills.
Preferred Skills
Experience with
Ansible playbooks and Ansible Automation Platform administration .
Experience with
Terraform , particularly Terraform Enterprise.
Knowledge of Docker and Kubernetes/OpenShift.
Experience with Git, Bitbucket, and CI/CD toolchains.
Knowledge of Agile methodologies.
Good understanding of JVMs and garbage collection mechanisms.
Experience with relational databases.
Application support, production release, and development team coordination experience.
Key Competencies
Strong analytical and problem-solving skills.
Ability to manage multiple priorities in high-pressure production environments.
Robust ownership and escalation-management capabilities.
Ability to automate repetitive operational processes and reduce system downtime.
Excellent stakeholder coordination and communication skills.
Willingness to participate in weekend and on-call support.
📌 Core SRE Engineer (Montreal)
🏢 AceStack
📍 Montreal