14 Sep
|
Jobtailor
|
Toronto
- Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems
- Partner with development teams as a reliability consultant and influence architectural decisions
- Write code to automate operational tasks and CI/CD pipelines
- Build internal tools, libraries, and frameworks for self-service observability
- Participate in a 24/7 on-call rotation and act as incident commander during critical disruptions
- Conduct blameless root cause analyses and implement corrective actions
- Monitor, measure, and optimize system performance, latency, and capacity
- Forecast capacity needs using usage patterns and historical data
- Build and integrate AIOps solutions, including automated responses and self-healing systems
- Use AI-assisted coding tools such as Claude Code and Cursor
- Develop and document runbooks and procedural guides for the observability knowledge base
- Analyze telemetry data, build predictive capacity models, and identify bottlenecks and failure modes
Requirements
- Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience
- 5+ years of professional experience in a Site Reliability Engineering, DevOps, or Software Engineering role focused on infrastructure and operations
- Robust programming proficiency in one or more high-level languages such as Rust, Go, Python, or Typescript
- Comfortable writing, testing,
and deploying production-grade code
- Deep knowledge of AWS services, especially networking, IAM, EKS, ALBs/NLBs, Route 53, and CloudWatch
- Proven experience with Kubernetes in production, including service exposure, networking, and availability engineering
- Solid understanding of Linux/Unix operating systems, TCP/IP, DNS, HTTP, and modern distributed systems architecture
Core Competencies
Demonstrates expertise in designing and maintaining scalable distributed systems, with a strong focus on automation, observability, and incident management. Proficient in programming and cloud services, particularly in AWS and Kubernetes, to optimize system performance and reliability.
Highest-signal resume keywords
- Site Reliability Engineering
- AWS Services
- Kubernetes
- Programming Proficiency
- Automation
Hard Skills
- Rust
- Go
- Python
- Typescript
- Linux/Unix
- TCP/IP
- DNS
- HTTP
- Distributed Systems Architecture
- CI/CD
Soft Skills
- Incident Management
- Root Cause Analysis
- Collaboration
Certifications & Qualifications
- Bachelor's Degree in Computer Science
Industry Keywords
- Infrastructure
- Operations
- Observability
- Capacity Forecasting
- Self-Healing Systems
Tools & Technologies
- AIOps Solutions
- Claude Code
- Cursor
- CloudWatch
- EKS
- ALBs/NLBs
- Route 53
📌 Senior Site Reliability Engineer, SRE (Toronto)
🏢 Jobtailor
📍 Toronto