30 Jul
|
Socket.dev
|
Toronto
30 Jul
Socket.dev
Toronto
Position: Senior SRE Engineer Job Type: Full Time Immediate Interview We are looking for a Senior Site Reliability Engineer (SRE) with a strong platform ownership mindset to drive reliability, scalability, and performance of mission-critical, distributed systems. This role sits at the intersection of software engineering, cloud infrastructure, and production operations, with a focus on building resilient systems, improving observability, automating operations, and driving reliability at scale. You will act as a technical lead for platform reliability, working closely with engineering and business stakeholders to ensure systems are highly available, performant, and continuously improving. 3+ years of experience in SRE, DevOps, or Production Engineering ~ Experience working in production-critical environments with high availability requirements ~ Own availability, performance, and scalability of production systems Drive continuous improvements in system resilience and efficiency Incident Management &
• Root Cause Analysis Perform deep root cause analysis across infrastructure, application, data, and network layers Implement long-term fixes and reduce recurrence through engineering improvements Observability &
• Monitoring Design and enhance monitoring, logging, and alerting systems Develop actionable dashboards and improve alert quality Automation &
• DevOps Practices Build and maintain CI/CD pipelines Troubleshoot distributed systems across compute, storage, and network layers Diagnose latency, routing,
and performance issues in globally distributed environments Data &
• Workflow Reliability Troubleshoot data pipelines, job failures, and data inconsistencies Perform data validation and analysis Ensure reliability across data dependencies and workflows Networking &
• Traffic Management Akamai or similar) to optimize traffic routing and performance Act as a liaison between engineering teams and business stakeholders Communicate system status, incidents, and risks with clarity and context AI-Driven Reliability (Emerging Focus) Apply AI/ML-driven techniques for anomaly detection, alert optimization, and Demonstrates strong ownership of production systems and outcomes Applies structured, analytical thinking to complex technical problems Communicates effectively in high-impact, production-critical scenarios Focuses on long-term reliability and scalability improvements Technical Skills: Programming &
• Automation Strong experience in Python for automation and tooling Hands-on experience with AWS, Azure, or GCP Strong understanding of cloud architecture, networking, and security fundamentals DevOps &
• CI/CD Strong understanding of build, release, and deployment pipelines Robust logging, monitoring, and alerting practices Data &
• Databases Strong SQL skills for troubleshooting and validation Understanding of data pipelines and system dependencies Strong Linux fundamentals Exposure to Kubernetes and web servers (e.g., Networking &
• CDN ~
📌 Senior Site Reliability Engineer/DevOps (Toronto)
🏢 Socket.dev
📍 Toronto