30 Sep
|
apptoza
|
Ontario
Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE) Toronto, On
Hybrid - 4 days
ABOUT THE ROLE We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.
This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.
WHAT YOU'LL DO OBSERVABILITY STACK OWNERSHIP Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents
Manage observability deployments using GitOps principles and infrastructure as code
Implement long-term metrics storage solutions with cloud object storage
Maintain and upgrade observability components across development, QA, UAT, production, and DR environments
Configure distributed observability architecture spanning multiple data centers and cloud providers
METRICS & MONITORING Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications
Create ServiceMonitors, PodMonitors for automated metrics collection
Develop rules for intelligent alerting with minimal false positives
Configure multi-cluster metrics federation and aggregation
Optimize metrics cardinality, storage e_iciency, and query performance
Implement recording rules for pre-aggregated metrics and SLI calculations
DASHBOARDS & VISUALIZATION Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics
Create reusable dashboard templates and libraries for development teams
Implement dashboard-as-code
Configure multiple datasources
Design executive dashboards with SLO/SLI tracking and business KPIs
Implement role-based access control and multi-tenancy in Grafana
LOGGING INFRASTRUCTURE Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or similar)
Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.)
Design log retention policies balancing cost, compliance, and operational needs
Create LogQL/Lucene queries and log-based alerts
Implement log correlation with metrics and traces for unified troubleshooting
Build log aggregation pipelines with parsing, filtering, and enrichment
ALERTING & INCIDENT MANAGEMENT Configure intelligent alerting with Alertmanager or equivalent platforms
Design alert rules with appropriate severity levels,
thresholds, and SLOs
Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks
Create automated runbooks and remediation workflows
Develop alert inhibition, silencing, and grouping strategies
Tune alerting to achieve signal-to-noise ratio improvements
Integrate with incident management and on-call rotation systems
DISTRIBUTED TRACING Implement distributed tracing solutions using OpenTelemetry, Jaeger, or Tempo
Instrument applications for trace collection and correlation
Configure trace sampling strategies for cost and performance optimization
Build trace-based dashboards for latency analysis and dependency mapping
Integrate tracing with metrics and logs for comprehensive observability
AI/ML FOR OBSERVABILITY Implement AI-powered anomaly detection for metrics and logs
Build predictive alerting using machine learning models to forecast issues before they occur
Develop intelligent alert correlation and root cause analysis systems
Integrate LLM-based tools for log analysis and troubleshooting assistance
Implement AIOps capabilities for automated incident triage and resolution
Use AI to optimize alert thresholds and reduce false positives
Build natural language query interfaces for observability data
Implement Model Context Protocol (MCP) for AI agent integration with observability platforms
AUTOMATION & PLATFORM INTEGRATION Automate observability deployment using GitOps workflows (FluxCD, ArgoCD)
Integrate with CI/CD pipelines for automated testing and validation
Build self-service portals for teams to create dashboards and alerts
Develop APIs and CLIs for observability automation
Integrate with secrets management solutions (Vault, AWS Secrets Manager)
Configure LDAP/AD/SSO authentication for observability platforms
Automate compliance reporting and audit logging
ENABLEMENT & COLLABORATION Onboard application teams to observability platforms
Provide guidance on instrumentation best practices
Create documentation, training materials, and self-service guides
Conduct workshops
Support development teams during incidents with observability insights
Collaborate with SRE, DevOps, and platform engineering teams
PERFORMANCE & COST OPTIMIZATION Monitor and optimize observability stack resource consumption
Implement autoscaling for stateless observability components
Tune data retention, compaction,
and downsampling strategies
Conduct capacity planning for metrics and log storage growth
Optimize query performance and dashboard response times
Implement cost allocation and chargeback for multi-tenant environments
REQUIRED QUALIFICATIONS EXPERIENCE 4-6 years of experience in observability, monitoring, SRE, or platform engineering roles
3+ years hands-on production experience with Prometheus and Grafana
2+ years working with Kubernetes and containerized environments
Robust expertise with PromQL for metrics querying and alerting
Experience deploying and managing observability stacks at scale (1000+ nodes)
Proven track record of reducing MTTR through e_ective observability
CORE TECHNICAL SKILLS OBSERVABILITY PLATFORMS: Prometheus (including Prometheus Operator)
Grafana (dashboards, alerting, plugins)
Thanos, Cortex, or Mimir for long-term storage
Alertmanager or equivalent alerting platforms
LOGGING: Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs
Log collection agents (Promtail, Fluentd, FluentBit, Vector)
LogQL, Lucene, or equivalent query languages
TRACING: OpenTelemetry (OTEL Collector, instrumentation)
Jaeger, Zipkin, Tempo, or AWS X-Ray
Trace sampling and correlation strategies
KUBERNETES & CONTAINERS: Kubernetes architecture and operations (1.24+)
Custom Resource Definitions (CRDs) and Operators
ServiceMonitor, PodMonitor, PrometheusRule resources
Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor)
GITOPS & AUTOMATION: Basic SQL for data analysis
CLOUD & INFRASTRUCTURE: AWS, Azure, or GCP cloud platforms
Multi-cloud and hybrid architectures
vSphere or on-premises virtualization (nice to have)
AI/ML & EMERGING TECHNOLOGIES Experience with AI/ML frameworks for observability (Prophet, TensorFlow, PyTorch)
Anomaly detection algorithms and time-series forecasting
LLM integration for log analysis and troubleshooting (GPT, Claude, etc.)
Model Context Protocol (MCP) for AI agent integration
AIOps platforms (Moogsoft, BigPanda, Datadog Watchdog, etc.)
Natural language processing for log parsing and analysis
Familiarity with vector databases for semantic search (Pinecone, Weaviate)
Experience with AI-powered root cause analysis tools
Knowledge of prompt engineering for observability use cases
PREFERRED QUALIFICATIONS Kubernetes certifications (CKA, CKAD, or CKS)
Grafana certification or equivalent training
Knowledge of service mesh observability (
Experience with SRE practices, SLIs, SLOs, and error budgets
Background in financial services or regulated industries
Familiarity with compliance requirements (SOX, PCI-DSS, etc.)
Contributions to open-source observability projects
Experience with APM tools
#J-18808-Ljbffr
📌 Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRA (Ontario)
🏢 apptoza
📍 Ontario