Elevate cloud performance as a Site Reliability Engineer with Insight Global in Vancouver. This hands-on role focuses on operational excellence, stability, and automation in engineering platforms. Join the Enterprise Platform Engineering team at Insight Global, where you will enhance reliability and performance while working closely with Lead Platform Engineers and GitLab experts.
Your expertise in observability and incident management will play a crucial role in creating robust cloud infrastructure and improving operational workflows. Key Responsibilities:
Design and maintain monitoring and observability solutions
Define and track Service Level Indicators and Objectives
Lead incident response and troubleshooting efforts
Automate operational processes to enhance reliability
Build platform health dashboards and operational reports Requirements:
5+ years in Site Reliability Engineering or related fields
Experience in AWS and Kubernetes settings
Proficient in tools like Grafana, Prometheus, and Terraform
Solid scripting skills in Python, Go, or Bash
Understanding of disaster recovery and scalability principles Drive reliability and automation in your engineering career with Insight Global!
📌 Site Reliability Engineer In Vancouver
🏢 Insight Global
📍 Vancouver
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.