12 Sep
|
Rootly
|
Toronto
Who you are
- In short, this role is designed for individuals that crave ownership, stimulating technical challenges, love shipping quick, and are mission driven
- 5+ years of experience in an SRE, Platform, or Infrastructure Engineering role
- 5+ years of experience writing software in a production environment
- Strong technical knowledge of cloud infrastructure, distributed systems, and reliability practices
- Strong understanding of observability, performance tuning, and scaling strategies
- Deep familiarity with incident response, monitoring, and CI/CD systems
- Hands-on experience supporting web or RPC services at meaningful scale
- You write code to solve infrastructure problems; not shell scripts alone, but production-grade software
- You have a big-picture systems mindset and a proactive approach to reliability
- You’ve embedded with product teams and influenced design and architecture decisions
- You’re comfortable taking ownership of complex problems—and seeing them through
- Experience with Ruby and Go is a plus
What the job involves
- This is an opportunity to join Rootly as an early SRE leader and shape our technical foundation. You will experience the balance of being scrappy and operating at scale
- What you’ll be doing one day could look very different the next. You will be empowered to identify opportunities that will help us grow and own it
- We won’t sugarcoat it, the work will be challenging, but it will also be one of the most rewarding learning experiences of your career
- Embed with product teams to enhance observability, reliability, and performance of their services
- Own our CI/CD pipelines, observability tooling, monitoring systems, and incident response processes
- Build tools and automation to eliminate manual toil, improve engineering velocity and developer experience, and improve system reliability
- Collaborate deeply across engineering to understand systems at the code level and surface cross-cutting reliability, performance, and scaling concerns
- Architect and scale our infrastructure, ensuring best-in-class performance, availability, and operational excellence
- Drive capacity planning efforts to ensure our infrastructure is resilient and scalable as we grow
- Define and manage SLOs and error budgets in partnership with Engineering teams who own production services
- Be vocal - act as a strong voice and force of reliability, quality, performance, and scalability
Benefits
- Medical, dental, vision, and life coverage for you and your dependents
- Flexible PTO, plus company holidays, team off-sites, and a year-end shutdown to recharge
- Meaningful equity in a high-growth company, and 401(k) matching for US employees
- Hybrid work (we hire across the US, Canada, Europe, and Singapore)
- Free lunches at HQ
- Fast-moving, high-impact environment where you have autonomy to own big problems from day one
- Top-tier hardware and tools, plus an annual stipends for home offices
📌 Senior Site Reliability Engineer (Toronto)
🏢 Rootly
📍 Toronto