10 Sep
|
Artech Information Systems
|
Toronto
10 Sep
Artech Information Systems
Toronto
Job Title: Sr Support Engineer
Location: Montreal, QC - Hybrid (2-4 Days WFO)
Duration: 06-12 Months
Hourly Pay Rate: 50 CAD/hr
Job Description:
Design and implement observability-as-code solutions using Terraform to deploy monitoring pipelines, dashboards, and alerting strategies across distributed systems. Drive observability improvements leveraging industry-leading tools to achieve real-time performance insights and comprehensive system visibility. Instrument applications for end-to-end observability, implementing distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures. Troubleshoot complex incidents in production environments, diagnosing root causes across multiple service layers, databases, caches, and APIs under load using SLI/SLO frameworks. Investigate and resolve Azure Kubernetes Service (AKS) infrastructure, ensuring reliability and scalability of containerized workloads with deep proficiency in Terraform and Azure managed services. Translate business requirements into observable, resilient systems that meet defined SLIs/SLOs and drive reliability improvements. Automate operational tasks to reduce toil and improve system resilience through infrastructure-as-code and CI/CD best practices. Lead incident response and remediation for mission-critical systems, conducting blameless postmortems and building resilience through chaos engineering and tabletop exercises. Collaborate cross-functionally with development, platform, and business teams to improve service availability, scalability, and operational excellence.
Required Skills & Qualifications (Must-have qualifications that candidates must meet to be considered)
- 8 years hands-on experience in observability, SRE,
or DevOps roles with proven expertise across infrastructure and application-level reliability.
- Deep expertise in observability tooling: Dynatrace, ELK, Splunk, and PagerDuty; demonstrated understanding of observability principles (instrumentation, correlation IDs, SLI/SLO frameworks).
- Advanced proficiency with Azure Kubernetes Service (AKS), Terraform, and Azure managed services; proven ability to design and implement infrastructure-as-code solutions.
- Solid hands-on experience instrumenting applications for comprehensive observability: distributed tracing, metrics collection, and log aggregation across Node.js and .NET applications.
- Proven troubleshooting expertise in distributed systems, diagnosing root causes across multiple service layers, databases, caches, and APIs in production environments.
- Excellent incident management skills: hands-on experience with PagerDuty and ServiceNow; ability to resolve high-severity incidents rapidly and conduct effective root cause analysis.
- Knowledge of incident, problem, and change management processes, including SRE principles, blameless postmortems, and chaos engineering practices.
- Exceptional communication and leadership skills.
Preferred Skills & Qualifications (Nice-to-have skills but are not required)
- Experience with other cloud platforms and services.
- Familiarity with additional programming languages and frameworks.
- Experience in financial services or a related industry.
Day-to-Day Responsibilities (key tasks and expectations for the role)
- Design and deploy observability solutions using Terraform.
- Drive improvements using industry-leading observability tools.
- Instrument applications for comprehensive observability.
- Troubleshoot and resolve complex production incidents.
- Collaborate with cross-functional teams to enhance service availability and reliability.
For immediate consideration please click APPLY to begin the screening process with Alex.
📌 Hiring Sr Support Engineer in Montreal, QC (Toronto)
🏢 Artech Information Systems
📍 Toronto