03 Aug
|
Releady
|
Toronto
OVERVIEWThis senior, client-facing observability role supports a major Washington-based airline. The engineer consults with the client’s internal engineering teams to identify systems, pain points, and reliability gaps, then designs and implements observability solutions—dashboards, metrics, SLIs/SLOs, alerting strategies, and visibility improvements. The role also helps define enterprise SRE standards and coaches teams to adopt consistent best practices.Duration: 12+ months (not C2C eligible)Location: Remote (PST Hours NOT in California)Rate: $50 - $62/hr DOEMust be able to work on W2 without sponsorshipRESPONSIBILITIESDay to dayMeet with internal teams to gather technical and operational requirementsDesign and implement tailored observability solutions across tools like Grafana, Sumo, AppDynamics, and New RelicBuild deeper dashboards for product teams and executive visibilityDefine and maintain SLOs, SLIs, and reliability reporting patternsIdentify gaps in monitoring or alerting and lead the solutioningPartner with embedded SREs across the Client’s hub and spoke modelInfluence tool consolidation, standards, and enterprise reliability strategyThis role acts as an internal consultant and technical leader for observability and reliability practices.Core Responsibilities:Build dashboards in Grafana for internal teams and leadership.Maintain observability tools and handle incoming requests.Connect data sources across tools (Grafana, Sumo, AppD, Current Relic).Assist teams with setting up alerting, logging structure,
and basic SLOs.Instrument new apps into monitoring tools.Create repeatable patterns and templates for team onboarding.Build playbooks and small automation tasks using Ansible Automation Platform.QUALIFICATIONS3+ years of hands-on observability experience (Grafana required plus supporting tools)2+ years practicing SRE fundamentals (SLOs/SLIs, incident patterns, distributed systems, reliability engineering)5+ total years in SRE, DevOps, cloud, systems, platform, or monitoring engineering rolesExperience partnering with application teams to gather requirements and deliver solutionsStrong ability to explain complex concepts clearly to non-SRE partnersRequired SkillsAdvanced Grafana expertise — design complex dashboards, build data transformations, define SLOs/SLIs, and integrate multiple data sources.SRE principles & systems thinking — deep knowledge of service health, SLOs/SLIs, error budgets, incident patterns, distributed systems, and reliability fundamentals.Cross-team collaboration & requirements gathering — engage with teams to understand needs, translate them into observability solutions, and deliver dashboards, alerting, and reliability patterns.PreferredExperience with ThousandEyes, AppDynamics, New Relic, or Sumo LogicFamiliarity with Azure, Kubernetes, CI and CD pipelines, or software delivery platformsExperience contributing to observability standards at scaleBackground in high uptime industries such as travel, finance, telecom, or cloud-based SaaSWe are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, disability status, or other non-merit factor. We are committed to creating a diverse and inclusive environment for all employees.#J-18808-Ljbffr
📌 Senior Sre Engineer - $50 - $62 An Hour - Remote (Toronto)
🏢 Releady
📍 Toronto