05 Oct
|
Staffingine
|
Toronto
05 Oct
Staffingine
Toronto
Job Title: Systems Reliability Engineer Job Location: Toronto, ON Job Type: Contract Job Description: Own the reliability, resiliency and availability of the Embedded Finance platform, proactively identifying and mitigating risks to service continuity.
Design, implement and maintain comprehensive monitoring and alerting frameworks leveraging Splunk, Dynatrace, Grafana and Datadog to provide end-to-end observability across the platform.
Define and track service level objectives (SLOs), service level indicators (SLIs) and error budgets to measure and improve platform health.
Lead and participate in incident response, serving as a technical driver during remediation calls and coordinating with impacted and impacting technical and product teams.
Own and advance the root cause analysis (RCA)
process - investigating incidents, documenting the sequence of events and remediating actions, and clearly identifying underlying root causes to prevent recurrence.
Ensure timely creation and management of incident tickets (e.g., Service
Now) and accurate incident tracking, aging and reporting.
Build automation and tooling to reduce toil, improve mean time to detection (MTTD) and mean time to resolution (MTTR), and increase operational efficiency.
Collaborate with engineering, product and risk stakeholders to embed reliability best practices into the platform lifecycle.
📌 Systems Reliability Engineer (Toronto)
🏢 Staffingine
📍 Toronto