Join our team as a Lead Site Reliability Engineer focused on enhancing system reliability and scalability. Drive incident management and automate workflows to improve operational efficiency. This role targets a technically skilled engineer committed to maintaining high availability for production systems.
You will collaborate closely with cross-functional teams to implement reliability practices and oversee incident management efforts. Build resilient systems while leveraging your cloud infrastructure expertise. Key Responsibilities:
- Drive continuous improvements in system resilience
- Lead root cause analysis and implement long-term fixes
- Build and maintain CI/CD pipelines
- Troubleshoot distributed systems effectively
- Communicate incidents clearly to stakeholders
Requirements:
- 3+ years of SRE, DevOps experience
- Robust SQL skills for validation and troubleshooting
- Hands-on experience with containerization technologies
- Familiarity with Infrastructure as Code approaches
- Knowledge of DNS and networking fundamentals
Make your mark on our mission-critical systems while embracing innovative reliability solutions.
📌 Lead Site Reliability Engineer Opening (Toronto)
🏢 Gemini Solutions
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.