04 Sep
|
AceStack
|
Toronto
SRE Technical Project Manager Location: Toronto, ON Work Arrangement: Onsite Employment Type: Full-Time FTE Job Summary: We are seeking an experienced SRE / Dev
Ops Engineer Technical Project Manager focused on reliability engineering, automation, observability, and cloud operations.
The ideal candidate will have strong hands-on expertise with Dynatrace, AWS, Azure, Ansible, Terraform, CI/CD, and Kubernetes , along with the ability to coordinate technical initiatives and drive reliability improvements.
Required Skills & Qualifications Strong expertise in Dynatrace , including APM, Davis AI, RUM, and infrastructure monitoring.
Experience using Dynatrace Davis AI for root cause analysis, anomaly detection, predictive insights, and alert optimization.
Hands-on experience with One
Agent, Smartscape, distributed tracing, SLIs/SLOs, dashboards, and alert management .
Strong automation experience using Ansible for deployments, provisioning, patching, and remediation.
Hands-on cloud experience with AWS (primary) and Azure , including serverless and cloud-native architectures.
Strong Infrastructure as Code (IaC) experience with Terraform, Cloud
Formation, or AWS CDK .
Experience implementing CI/CD pipelines using Jenkins, Git
Hub Actions, and Git
Lab CI .
Experience integrating monitoring and observability tools with CI/CD pipelines.
Strong knowledge of Docker, Kubernetes, ECS, and AKS .
Experience with High Availability (HA), Disaster Recovery (DR), incident response, and reliability engineering .
Robust programming/scripting skills in Python (boto3) and Bash .
Experience with AWS Cloud
Watch and Azure Monitor .
Exposure to Prometheus, Grafana, and ELK is an advantage.
Key Responsibilities Design, implement,
and maintain reliable, scalable, and highly available cloud infrastructure.
Lead observability initiatives using Dynatrace across applications, infrastructure, and user experience.
Configure and optimize Dynatrace One
Agent, Smartscape, distributed tracing, dashboards, alerts, SLIs, and SLOs.
Leverage Davis AI for automated anomaly detection, root cause analysis, predictive insights, and alert optimization.
Develop and maintain Ansible playbooks for deployment, provisioning, patching, and automated remediation.
Automate infrastructure provisioning and configuration using Terraform, Cloud
Formation, or CDK .
Build and maintain CI/CD pipelines and integrate observability and monitoring capabilities into deployment workflows.
Support containerized workloads using Docker, Kubernetes, ECS, and AKS .
Implement and maintain HA, DR, monitoring, alerting, and incident response processes.
Troubleshoot complex application, infrastructure, cloud, and network reliability issues.
Collaborate with development, infrastructure, security, cloud, and business teams to improve system reliability and operational efficiency.
Drive automation and continuous improvement across SRE and Dev
Ops processes.
Provide technical leadership and coordinate delivery of reliability, observability, and automation initiatives.
Preferred Qualifications Experience in Site Reliability Engineering (SRE), Dev
Ops, Cloud Engineering, or Technical Project Management .
Strong understanding of cloud-native architecture and enterprise observability.
Excellent communication, stakeholder management, problem-solving, and technical leadership skills.
Experience managing multiple technical initiatives in a fast-paced enterprise environment.
📌 SRE – Technical Project Manager || Toronto, ON - Onsite || Fulltime FTE
🏢 AceStack
📍 Toronto