29 Sep
|
AceStack
|
Halifax Regional Municipality
29 Sep
AceStack
Halifax Regional Municipality
OBSERVABILITY ENGINEER
Job Type: Full Time
Location: Halifax, NS
Work Model: Onsite / 4 Days
Primary Skills: Prometheus & Grafana, Kubernetes Observability, GitOps & Infrastructure as Code Job Description
We are seeking an experienced OBSERVABILITY ENGINEER to join an Enterprise Kubernetes Platform team within a leading financial services organization. The role will own the observability stack across 50+ production Kubernetes clusters, delivering metrics, logging, tracing, and alerting capabilities for mission-critical applications.
The position combines deep expertise in contemporary observability platforms with emerging AI/ML capabilities to support intelligent monitoring, predictive alerting, and self-healing infrastructure. Key Responsibilities
- Design, deploy, and maintain enterprise-scale observability infrastructure using Prometheus, Grafana, Thanos, Loki, and modern collection agents.
- Manage observability deployments using GitOps principles and Infrastructure as Code.
- Implement long-term metrics storage using cloud object storage.
- Maintain and upgrade observability components across Development, QA, UAT, Production, and DR environments.
- Configure distributed observability architectures across multiple data centers and cloud providers.
- Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications.
- Create ServiceMonitors and PodMonitors for automated metrics collection.
- Develop intelligent alerting rules with minimal false positives.
- Configure multi-cluster metrics federation and aggregation.
- Optimize metrics cardinality, storage efficiency, and query performance.
- Support scalable monitoring and observability solutions for enterprise Kubernetes environments.
- Explore AI/ML capabilities for predictive monitoring, intelligent alerting, and automated remediation.
Required Skills & Qualifications
- Strong hands-on experience with Prometheus, Grafana, Thanos, and Loki.
- Extensive experience with Kubernetes monitoring and observability.
- Experience with GitOps methodologies and Infrastructure as Code.
- Strong understanding of metrics collection, logging, tracing, alerting, and monitoring architecture.
- Experience with ServiceMonitors, PodMonitors, metrics federation, and multi-cluster monitoring.
- Knowledge of cloud object storage and distributed observability architectures.
- Experience supporting production, QA, UAT, and DR environments.
- Solid troubleshooting and performance optimization skills.
- Understanding of AI/ML-driven monitoring, predictive alerting, or self-healing infrastructure is preferred.
📌 OBSERVABILITY ENGINEER (Halifax Regional Municipality)
🏢 AceStack
📍 Halifax Regional Municipality