Observability Engineer Kubernetes, Prometheus, Grafana & Cloud Monitoring Role Overview Experienced Observability Engineer to join an Enterprise Kubernetes Platform team within a leading financial services organization Own and manage the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities Build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure using modern observability tools and AI/ML capabilities Key Responsibilities Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and contemporary collection agents Manage observability deployments using GitOps principles and Infrastructure as Code Implement long-term metrics storage solutions using cloud object storage Maintain and upgrade observability components across development, QA, UAT, production,
and DR environments Configure distributed observability architecture spanning multiple datacenters and cloud providers Metrics & Monitoring Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications Create Service Monitors and Pod Monitors for automated metrics collection Develop intelligent alerting rules with minimal false positives Configure multi-cluster metrics federation and aggregation Optimize metrics cardinality, storage efficiency, and query performance Essential Skills Kubernetes Observability Prometheus Grafana Thanos Loki Metrics, Logging, and Tracing Alerting and Monitoring GitOps Infrastructure as Code Cloud Monitoring Kubernetes Platform Engineering AI/ML-Based Monitoring Solutions
5+
Sailpoint
Required Skill Profession
Other General
📌 Observability Engineer (Toronto)
🏢 Astra North Infoteck
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.