31 Aug
|
Artech Information Systems
|
Toronto
31 Aug
Artech Information Systems
Toronto
Title: Site Reliability Engineer (SRE) Business Analyst
Location: Toronto, ON Hybrid (2 days per week in-person at Toronto office preferred)
Duration: 6 Months
Pay Range :C$49 INC
Skills Required: Digital : DevOps~Digital : Site Reliability Engineering (SRE)~Dynatrace
Years Experience: 8-10
Role Summary:
We are seeking a highly skilled SRE / DevOps Engineer with strong expertise in Dynatrace, AI-driven observability (Davis AI), automation (Ansible), and cloud platforms (AWS & Azure). This role will focus on proactive monitoring, intelligent automation, and reliability engineering, ensuring high system availability and performance across distributed environments.
Key Responsibilities:
1. Dynatrace & AI-Driven Observability (Primary Focus)
Lead implementation and optimization of Dynatrace platform across applications and infrastructure
Leverage Dynatrace Davis AI for:
Automated root cause analysis
Anomaly detection and event correlation
Predictive performance insights
Alert noise reduction
Configure and manage:
OneAgent deployments
Smartscape topology mapping
Service flow and distributed tracing
Define and monitor SLIs, SLOs, and user experience metrics
Build custom dashboards, alerts, and observability pipelines
Integrate Dynatrace with:
CI/CD pipelines (release validation, performance gating)
Incident management tools (PagerDuty, ServiceNow, etc.)
Enable self-healing automation using Dynatrace event triggers and AI insights
2. Automation & Configuration Management (Ansible Focus)
Design and implement automation using Ansible for:
Configuration management
Application deployments
Setting provisioning
Develop reusable playbooks and roles for scalable operations
Automate operational tasks, patching, and compliance processes
Integrate Ansible with CI/CD pipelines and monitoring systems
Improve system reliability through automated remediation workflows
3. Cloud & AWS DevOps Tooling
Design and manage cloud-native systems on AWS, with exposure to Azure
Develop infrastructure using Terraform, CloudFormation, or CDK
Build and manage CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI)
Develop and deploy serverless architectures (Lambda, API Gateway, Step Functions)
Use AWS SDK (boto3) to automate DevOps and operational workflows
Deploy and maintain large-scale production systems via automated pipelines
Optimize cloud infrastructure for cost, performance, and scalability
4. Monitoring, Logging & Multi-Cloud Observability
Strong experience with:
AWS CloudWatch (metrics, logs, alarms, dashboards)
Azure Monitor / Log Analytics
Design unified observability across multi-cloud environments
Implement logging and tracing strategies for distributed systems
5. Containers, Platforms & Reliability Engineering
Work in containerized environments (Docker)
Manage orchestration platforms such as Kubernetes, ECS, AKS
Ensure high availability using:
Fault tolerance design
Disaster recovery strategies
Support incident response, on-call processes, and root cause analysis (RCA)
Required Qualifications:
Proven experience with Dynatrace (APM, RUM, infrastructure monitoring)
Strong hands-on experience with Dynatrace Davis AI capabilities
Experience with Ansible for automation and configuration management
Deep knowledge of AWS services and cloud-native architectures
Experience with Infrastructure as Code tools (Terraform/CloudFormation/CDK)
Proficiency in Python (boto3), Bash scripting
Experience working in production-scale environments
Business Analyst experience
Scrum Master experience
Nice to Have:
Dynatrace certification (Associate/Professional)
Advanced experience with Dynatrace APIs and automation
Experience building self-healing systems using AI-driven triggers
Familiarity with Prometheus, Grafana, ELK stack
Azure cloud experience and certifications
Experience with GitOps and platform engineering
Comments for Suppliers:
📌 Site Reliability Engineer (SRE) Business Analyst (Toronto)
🏢 Artech Information Systems
📍 Toronto