05 Sep
|
Socket.dev
|
Toronto
05 Sep
Socket.dev
Toronto
Job Description WHAT IS THE PROSPECT? This role is responsible for designing, implementing, and maintaining SRE (Site Reliability Engineering) and AIOps (Artificial Intelligence for IT Operations) capabilities to ensure system reliability, proactive monitoring, and automation of self-healing operations. In addition to day-to-day support, the position provides end-to-end operational ownership across systems managed by multiple enterprise teams, including incident coordination, dependency management, and escalation. The role is also responsible for key security and compliance functions such as service ID and certificate management, SSO updates, vulnerability remediation, and lifecycle management of end-of-life components. Our team supports a portfolio of multi-platform HR data pipelines that move and process data into Snowflake through multiple integrated components, requiring end-to-end monitoring, coordination, and support across systems.In addition, we support SaaS-based applications that are primarily vendor-managed, while we retain responsibility for integration, access management, monitoring, and operational oversight. WHAT WILL YOU DO? Key Responsibilities: SRE & Reliability Engineering Define and operationalize SLIs, SLOs, and error budgets Own the incident management lifecycle (detection → triage → resolution → RCA → prevention) Lead problem management and eliminate recurring issues Develop and maintain runbooks, playbooks,
and recovery procedures Drive resilience engineering , including failover testing and capacity planning AIOps, Observability & Logging Implement and optimize AIOps capabilities using platforms such as Moogsoft Leverage Dynatrace for deep APM insights Integrate alerting and escalation workflows with PagerDuty Utilize synthetic monitoring via Catchpoint Design and maintain centralized logging solutions using the ELK stack (Elasticsearch, Logstash, Kibana) and enterprise Logging as a Service (LaaS) platforms Perform log analysis, correlation, and anomaly detection to support proactive issue identification Drive event noise reduction and intelligent alerting strategies Build dashboards and observability KPIs for operational insights Automation & Self-Healing Systems Design and implement automation-first solutions using Ansible and scripting (Python, Bash) Enable self-healing capabilities (auto-remediation, restart logic, workflow recovery) Orchestrate workflows using Stonebranch Reduce operational toil through automation and continuous improvement Application & Data Platform Support Provide L2/L3 support for data pipelines and integration workflows , including Snowflake ingestion and transformation processes Support Snowflake pipelines , ETL workflows, and orchestration dependencies Troubleshoot across distributed systems including: Object storage (e.G., S3) APIs, messaging, and file
📌 Senior Sre/Aiops Engineer (Toronto)
🏢 Socket.dev
📍 Toronto