Be a pivotal Site Reliability Engineer focused on improving infrastructure resilience and reliability. Collaborate remotely to drive operational success and enhance system performance in a energetic setting. This role allows you to design and operate reliable systems, playing a key role in preventing incidents and responding effectively.
By leading initiatives for continuous improvement and managing operational standards, you will ensure our services are robust and capable of scaling as necessary.
Key Responsibilities:
Drive improvements in system reliability and scalability
Oversee incident management and automation
Provide support via on-call rotations for critical services
Define SLIs, SLOs, and error budgets for decision-making
Enhance monitoring and observability across systems Requirements:
Experience with cloud environments like AWS
Solid knowledge of chaos engineering tools
Proven ability to troubleshoot live production systems
Familiarity with scripting and programming practices
Self-starter attitude in quick-paced work settings Utilize your skills to streamline operations and enhance system reliability, ensuring a solid foundation for ongoing technological advancements.
J-18808-Ljbffr
📌 Site Reliability Engineer For Cloud Infrastructure Management London (Canada)
🏢 Newton
📍 Canada