Elevate platform performance as a Senior Site Reliability Engineer in Toronto. Leverage your expertise in cloud infrastructure and automation to enhance system reliability and scalability. We are seeking a Senior Site Reliability Engineer to own and optimize production systems in a full time role.
You will work with large-scale distributed systems, ensuring high availability and performance while collaborating with diverse engineering teams. Your technical leadership will focus on incident management, observability, and continuous system improvements. Key Responsibilities:
Own the scalability and reliability of production systems
Lead end-to-end incident response and service restoration
Design and enhance monitoring and alerting systems
Automate operational workflows to minimize manual tasks
Troubleshoot latency and performance issues in distributed settings Requirements:
3+ years in SRE, DevOps, or Production Engineering
Experience with AWS, Azure, or GCP cloud platforms
Proficiency in Python and shell scripting
Familiarity with CI/CD tools like Jenkins
Solid SQL skills for troubleshooting and data validation Utilize your expertise in SRE practices to drive system reliability and performance at scale.
📌 Senior Site Reliability Engineer Toronto
🏢 Socket.dev
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.