Site Reliability Engineer (Sre) – Observability (Toronto)

Site Reliability Engineer (Sre) – Observability (Toronto)

10 Sep
|
Confidential
|
Toronto

10 Sep

Confidential

Toronto

Job Description: Site Reliability Engineer (SRE) – ObservabilityToronto - Hybrid (1-2 days office)We are looking for a Observability Engineer to help implement, operate, and improve observability capabilities across our applications and platforms. This role focuses on hands‑on onboarding, instrumentation, dashboarding, and alerting, working under established standards and guidance from senior engineers.You will collaborate with application, SRE, and operations teams to ensure systems are observable, supportable, and production‑ready.Key ResponsibilitiesObservability ImplementationImplement and maintain metrics, logs, and traces for applications and infrastructureAssist with onboarding applications into observability platforms (e.G., Dynatrace, ELK, Datadog)Configure dashboards, alerts, and basic anomaly detectionApplication Support & InstrumentationWork with development teams to enable structured logging, basic distributed tracing, and core metricsValidate observability requirements during Production Readiness Reviews (PRR)Troubleshoot missing or low‑quality telemetryMonitoring & AlertingConfigure alerts based on golden signals (latency, errors, traffic, saturation)Help reduce alert noise by tuning thresholds and alert logicSupport incident response by gathering logs, metrics,



and tracesOperations & ReliabilitySupport root cause analysis using observability toolsMaintain dashboards and documentation used by on‑call and support teamsParticipate in on‑call rotations (as applicable)Automation & Continuous ImprovementAssist in automating observability onboarding and validation tasksCreate and maintain reusable dashboards and alert templatesFollow established observability standards and best practicesRequired Qualifications2–4 years of experience in Observability, or SREWorking knowledge of metrics, logs, and basic tracing conceptsHands‑on experience with at least one observability platform (Dynatrace, Elastic/ELK, Datadog, Recent Relic, etc.)Basic understanding of SLIs/SLOs and service health indicatorsExperience with cloud platforms or hybrid environmentsAbility to write scripts (Python, Bash, PowerShell) for automation and troubleshootingPreferred QualificationsExperience with OpenTelemetry or APM agentsFamiliarity with Kubernetes or containerized workloadsExperience working with incident management tools (PagerDuty, ServiceNow)Exposure to Dynatrace/Kibana ELK or similar cloud‑native monitoringExperience in regulated or enterprise environments#J-18808-Ljbffr

📌 Site Reliability Engineer (Sre) – Observability (Toronto)
🏢 Confidential
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) – observability (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) – observability (toronto) / toronto