Maximize cloud database reliability as a Senior SRE at Grafana Labs. This fully remote role involves strategic automation and customer incident management.
We are on the lookout for a Senior Software Engineer specialized in Site Reliability Engineering. You will be essential in enhancing the reliability of our cloud databases while collaborating with product teams to address complex customer environments efficiently. This position focuses on incident response, service-level objective development, and proactive reliability practices in a rapid-paced setting.
Key Responsibilities:
• Collaborate intensely with engineering squads for optimization
• Ensure reliability for multi-tenant cloud databases
• Create and implement effective automation solutions
• Conduct reviews for incident responses and root causes
• Elevate observability and alerting mechanisms
Requirements:
• Over 6 years of engineering experience, minimum 3 in SRE
• Profound knowledge of Kubernetes and cloud infrastructure
• Expertise in programming languages such as Go or Python
• Strong incident response experience with blame-free culture
• Self-directed ability to collaborate in an engineering environment
Enhance operational resilience in cloud services while growing with Grafana Labs.
#J-18808-Ljbffr