Step into a pivotal role as a Senior Site Reliability Engineer, enhancing cloud infrastructure reliability and performance. Focus on automation, observability, and incident management to elevate our platform.
We are looking for a proactive Senior Site Reliability Engineer passionate about system integrity. Collaborate with engineers and business stakeholders to drive improvements in mission-critical environments. Optimize cloud infrastructure by focusing on scalability and system reliability.
Key Responsibilities:
• Own and monitor availability and performance of systems
• Implement SLIs, SLOs, and error budgets effectively
• Design monitoring and alerting systems for proactive issue detection
• Automate operational workflows and CI/CD processes
• Diagnose performance issues in distributed environments
Requirements:
• Minimum 3 years of experience in SRE or related fields
• Expertise in cloud platforms like AWS or GCP
• Robust Python scripting abilities for automation
• Experience with observability tools such as Grafana
• Understanding of CI/CD tools and practices
Drive reliability and performance improvements while leveraging your expertise in cloud-based solutions.
#J-18808-Ljbffr