05 Sep
|
Priceline
|
Toronto
Role OverviewThis role is eligible for our hybrid work model: Two days in-office. As a Site Reliability Engineer – Observability, you will play a key part in maturing our observability capabilities by standardizing instrumentation, improving telemetry quality, and enabling faster root cause analysis that directly impacts MTTR and MTTD.Responsibilities- Support and evolve end-to-end observability solutions for collecting, shipping, storing, and querying OpenTelemetry signals (metrics, logs, and traces) across infrastructure, containers, and Kubernetes environments.- Administer and operate core observability platforms (Splunk, New Relic, ClickHouse, Grafana, Lightrun), including service onboarding, access management, configuration, upgrades, and ongoing platform health.- Contribute to building and advancing a modern OpenTelemetry-based observability ecosystem that supports multiple telemetry types at scale.- Improve and standardize instrumentation practices across services, driving consistent logging, metrics, and distributed tracing implementation.- Partner with product and platform engineering teams to enhance production visibility and support SLO-driven reliability practices.- Optimize telemetry pipelines for performance, data quality, scalability, and cost efficiency.- Help define and support governance standards for observability, ensuring consistency, reliability, and scalability across teams.- Contribute to evolving our observability platform toward intelligent and AI-enabled capabilities, exploring opportunities to integrate AI or MCP-based solutions to improve signal quality, incident triage,
and operational efficiency.- Ensure observability platform reliability, security, and performance meet defined SLAs and operational standards.Required Qualifications- Bachelor’s degree in Computer Science or equivalent practical experience.- 3+ years of experience in Observability, SRE, DevOps, or platform engineering roles supporting production systems.- Strong understanding of APM and SRE fundamentals, including MELT (Metrics, Events, Logs, Traces), latency analysis, error rate monitoring, service dependency mapping, SLIs/SLOs, alert tuning, and root cause analysis.- Hands‑on experience administering at least one modern observability/APM platform (e.G., Splunk, New Relic, Grafana). Practical exposure to metrics, logs, distributed tracing, and platform configuration.- Experience supporting full-stack observability coverage across infrastructure, application, and browser monitoring layers.- Experience building dashboards and actionable alerts, including configuring alert workflows and integrations with incident management tools such as PagerDuty.- Experience implementing or supporting OpenTelemetry-based instrumentation and improving telemetry quality across services.- Familiarity with Kubernetes and cloud‑native environments – understanding of how applications are deployed, monitored,
and scaled.- Experience managing telemetry pipelines and agents (e.G., collectors, forwarders, sidecars), including onboarding services and troubleshooting ingestion issues.- Working knowledge of scripting or automation (e.G., Shell, Python) and CI/CD concepts.- Experience or familiarity with infrastructure‑as‑code tools such as Terraform for managing platform configurations and integrations is a plus.- Comfortable collaborating with engineering teams to improve monitoring standards, instrumentation quality, and overall production visibility.- Relevant certifications such as New Relic APM Practitioner, Reliability Engineer – Professional, Splunk Admin, or GCP Associate Cloud Engineer are a plus.- Demonstrated history of living the values important to Priceline: Customer, Innovation, Team, Accountability, and Trust.SalarySalary range: $110,000–$130,000 CAD per year, plus potential bonus and/or equity.Benefits- Health and wellness coverage (medical, dental, vision, mental health resources)- Generous paid time off, holidays, and Priceline Pause reset week- Work/life support (remote options, parental leave, dependent care)- Financial security programs (retirement plans, life and disability coverage, tax‑advantaged accounts)- Travel perks (employee‑only discounts, VIP deals, Big Deal Bucks credits)- Additional perks (partner discounts, tuition support, legal support, pet perks)Priceline is a proud equal opportunity employer. We embrace and celebrate the unique lenses through which our employees see the world.#J-18808-Ljbffr
📌 Site Reliability Engineer, Observability - C$110,000 - C$130,000 A Year (Toronto)
🏢 Priceline
📍 Toronto