- Monitoring / Observability tools - Dynatrace, ELK etc.
- Platform/ cloud Observability - OpenShift, Prometheus / Azure Cloud etc.
Key Responsibilities:
- Collaborate with various Infrastructure, Applications, platforms, and cloud teams on Observability solutions.
- Implement monitoring solutions using APM tools and Grafana for visualization - setup, configuration and developing monitoring / alerting solutions.
- Manage Grafana platform with team-specific dashboards covering various KPIs & data sources, enable with alerts and establish SLOs.
- Troubleshoot and resolve issues related to Observability solutions - Gaps, challenges and addressing solutions part of Production incidents.
- Analyze Infrastructure systems, services, and technologies towards monitoring, alerting and Incident response needs.
- Work in apps, platforms and infra services on resilient infrastructure, scalable, and highly available setting.
- Collaborate with App and services teams/SMEs to integrate monitoring solutions through Automation - APIs, webhooks, CI/CD deployments.
- Document system configurations, standard operating procedures, and best practices.
- Reflect on latest technologies and trends in Enterprise technologies, platforms, Automation and AI based solutions.
Seniority level
Mid-Senior level
Employment type
Contract
Job function
Other
Industries
IT Services and IT Consulting
#J-18808-Ljbffr
📌 Site Reliability Engineer (Ottawa)
🏢 Vertex Elite
📍 Ottawa
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.