17 Aug
|
TEEMA
|
Vancouver
What you will be doing: Provide technical and operational support for customers according to defined SLAs.
Operate, maintain, and support Kubernetes-based systems (EKS) within cloud-based AWS environments.
Design, build, and maintain advanced automation systems to enhance reliability, monitoring, and operational efficiency across production environments.
Develop scalable monitoring and alerting solutions to proactively detect issues before they impact customers.
Build and maintain comprehensive runbooks for NOC/SOC teams.
Serve as Tier-2 escalation for production incidents, collaborating with Dev
Ops teams and participating in a 247 on-call rotation.
Implement automation-driven improvements using scripts and configuration management tools to streamline operations.
Work closely with the Cloud Dev
Ops team to transition products from development to production via continuous integration and deployment processes.
What you must have: Experience:At least 3 years of experience in an SRE, Dev
Ops,
or similar cloud/monitoring support role in production environments.
Cloud & Containers: Hands-on experience with AWS infrastructure and solid practical knowledge of Kubernetes (EKS).
Systems & Scripting: Proficient in Linux system administration and scripting with Bash to automate operational processes.
Operations & Observability: Direct experience working with cloud monitoring, management, and alerting tools, alongside strong incident troubleshooting skills.
Flexibility & Soft Skills: Willingness to participate in a 247 on-call rotation, assertive and fast learner, with solid interpersonal and collaborative Salary/Rate: $70.00/hour Thank you for your interest in this opportunity.
If you are selected to move forward in the process, we will contact you directly.
If you do not hear from us, we encourage you to continue visiting our website for other roles that may be a good fit.
📌 Site Reliability Engineer (SRE) (Vancouver)
🏢 TEEMA
📍 Vancouver