We are seeking a Site Reliability Engineer (SRE) with strong expertise in observability, monitoring, and distributed tracing to join our SRE team. The ideal candidate will help us design, build, and scale an observability framework that provides end-to-end visibility into our systems and applications.
Responsibilities
- Design, implement, and maintain observability solutions
- Define and maintain SLIs, SLOs, and error budgets to measure and improve system reliability.
- Partner with development and operations teams to instrument applications and services for better monitoring and tracing coverage.
- Develop dashboards, alerts, and visualizations to provide actionable insights into system health and performance.
- Contribute to automation and self-healing practices that improve uptime and reduce operational toil.
- Stay current with trends in observability and advocate best practices across the engineering organization.
Requirements
- 7+ years of SRE/ Devops/ Cloud/ Infrastructure engineering experience with a focus on monitoring and observability.
- Strong communication skills with the ability to articulate technical requirements and explain perks of observability clearly to development teams.
- Proficiency with observability stacks such as Prometheus, Grafana, Loki, Elastic Stack, or Splunk Observability (Splunk/AppDynamics).
- Strong knowledge on cloud platforms (Google Cloud Platform, or Azure/AWS).
- Hand-on experience on container orchestration using Kubernetes (OCP, GKE, AKS)
- Familiarity with CI/CD pipelines like (Jenkins and Github actions), infrastructure as code (Terraform/Ansible/ARM/CloudFormation).
- Experience provisioning infrastructure and capacity planning.
📌 Site Reliability Engineer (Toronto)
🏢 Envision Technology Solutions
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.