Observability Engineer (Toronto)

Observability Engineer (Toronto)

03 Oct
|
apptoza
|
Toronto

03 Oct

apptoza

Toronto

Role: Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE) Location: Toronto, ON – 4 days onsite Contract Role ABOUT THE ROLE

We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.

This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.

OBSERVABILITY STACK OWNERSHIP

Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents

Manage observability deployments using GitOps principles and infrastructure as code

Implement long-term metrics storage solutions with cloud object storage

Maintain and upgrade observability components across development, QA, UAT, production, and DR environments

Configure distributed observability architecture spanning multiple data centers and cloud providers

METRICS & MONITORING

Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications

Create ServiceMonitors, PodMonitors for automated metrics collection

Develop rules for intelligent alerting with minimal false positives

Configure multi-cluster metrics federation and aggregation

Optimize metrics cardinality, storage efficiency, and query performance

Implement recording rules for pre-aggregated metrics and SLI calculations

DASHBOARDS & VISUALIZATION

Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics

Create reusable dashboard templates and libraries for development teams

Implement dashboard-as-code

Configure multiple datasources

Design executive dashboards with SLO/SLI tracking and business KPIs

Implement role-based access control and multi-tenancy in Grafana

LOGGING INFRASTRUCTURE

Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or Similar)

Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.)

Design log retention policies balancing cost, compliance, and operational needs





Create LogQL/Lucene queries and log-based alerts

Implement log correlation with metrics and traces for unified troubleshooting

Build log aggregation pipelines with parsing, filtering, and enrichment

ALERTING & INCIDENT MANAGEMENT

Configure intelligent alerting with Alertmanager or equivalent platforms

Design alert rules with appropriate severity levels, thresholds, and SLOs

Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks

Create automated runbooks and remediation workflows

Develop alert inhibition, silencing, and grouping strategies

Tune alerting to achieve signal-to-noise ratio improvements

Integrate with incident management and on-call rotation systems

DISTRIBUTED TRACING

Implement distributed tracing solutions using OpenTelemetry, Jaeger, or Tempo

Instrument applications for trace collection and correlation

Configure trace sampling strategies for cost and performance optimization

Build trace-based dashboards for latency analysis and dependency mapping

Integrate tracing with metrics and logs for comprehensive observability

AI/ML FOR OBSERVABILITY

Implement AI-powered anomaly detection for metrics and logs

Build predictive alerting using machine learning models to forecast issues before they occur

Develop intelligent alert correlation and root cause analysis systems

Integrate LLM-based tools for log analysis and troubleshooting assistance

Implement AIOps capabilities for automated incident triage and resolution

Use AI to optimize alert thresholds and reduce false positives

Build natural language query interfaces for observability data

Implement Model Context Protocol (MCP) for AI agent integration with observability platforms

AUTOMATION & PLATFORM INTEGRATION

Automate observability deployment using GitOps workflows (FluxCD, ArgoCD)

Integrate with CI/CD pipelines for automated testing and validation

Build self-service portals for teams to create dashboards and alerts

Develop APIs and CLIs for observability automation

Integrate with secrets management solutions (Vault,



AWS Secrets Manager)

Configure LDAP/AD/SSO authentication for observability platforms

Automate compliance reporting and audit logging

ENABLEMENT & COLLABORATION

Onboard application teams to observability platforms

Provide guidance on instrumentation best practices

Create documentation, training materials, and self-service guides

Conduct workshops

Support development teams during incidents with observability insights

Collaborate with SRE, DevOps, and platform engineering teams

PERFORMANCE & COST OPTIMIZATION

Monitor and optimize observability stack resource consumption

Implement autoscaling for stateless observability components

Tune data retention, compaction, and downsampling strategies

Conduct capacity planning for metrics and log storage growth

Optimize query performance and dashboard response times

Implement cost allocation and chargeback for multi-tenant environments

REQUIRED QUALIFICATIONS EXPERIENCE

4-6 years of experience in observability, monitoring, SRE, or platform engineering roles

3+ years hands-on production experience with Prometheus and Grafana

2+ years working with Kubernetes and containerized environments

Solid expertise with PromQL for metrics querying and alerting

Experience deploying and managing observability stacks at scale (1000+ nodes)

Proven track record of reducing MTTR through effective observability

CORE TECHNICAL SKILLS OBSERVABILITY PLATFORMS:

Prometheus (including Prometheus Operator)

Grafana (dashboards, alerting, plugins)

Thanos, Cortex, or Mimir for long-term storage

Alertmanager or equivalent alerting platforms

LOGGING

Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs

Log collection agents (Promtail, Fluentd, FluentBit, Vector)

LogQL, Lucene, or equivalent query languages

TRACING

OpenTelemetry (OTEL Collector, instrumentation)

Jaeger, Zipkin, Tempo, or AWS X-Ray

Trace sampling and correlation strategies

KUBERNETES & CONTAINERS:

Kubernetes architecture and operations (1.24+)

Custom Resource Definitions (CRDs) and Operators

ServiceMonitor, PodMonitor, PrometheusRule resources

Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor)

GITOPS & AUTOMATION:

Basic SQL for data analysis

CLOUD & INFRASTRUCTURE:

AWS, Azure, or GCP cloud platforms

Multi-cloud and hybrid architectures vSphere or on-premises virtualization (nice to have)

AI/ML & EMERGING TECHNOLOGIES

📌 Observability Engineer (Toronto)
🏢 apptoza
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: observability engineer (toronto) / toronto