- Monitoring and Alerting: Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users.
- Incident Response: Respond to incidents and outages, diagnose problems and implement solutions to minimize downtime and restore service.
- Automation: Automate repetitive tasks and processes to improve efficiency and reduce manual effort.
- Performance Optimization: Identify and address performance bottlenecks to ensure systems run efficiently and effectively.
- Infrastructure Management: Manage and maintain the underlying infrastructure including servers, networks, and cloud resources.
- Capacity Planning: Plan for future capacity needs to ensure systems can handle anticipated workloads.
- Release Engineering:
Develop and maintain processes for deploying software updates and releases.
- Collaboration: Work closely with developers, operations teams and other stakeholders to ensure system reliability and availability.
- Documentation: Maintain transparent and concise documentation of systems, processes and procedures.
- Continuous Improvement: Identify areas for improvement and implement changes to enhance system reliability and performance.
- Skills and Qualifications: Cloud Platform Microsoft Azure.
Desired Skills
- Excellent knowledge of AKS
- Monitoring tools: Dynatrace, Splunk, Grafana
- Operating Systems: Windows, Linux
- Scripting: Shell Scripting, Python, PowerShell
- Database: MySQL, Oracle, SQL database management
- Container Services: Kubernetes, Docker, Helm
- Understanding of Camunda is preferable
Seniority Level
Mid-Senior level
Employment Type
Contract
Job Function
Engineering and Information Technology
Industry
IT Services and IT Consulting
#J-18808-Ljbffr
📌 Azure SRE Developer (Toronto)
🏢 Aarorn Technologies
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.