04 Aug
|
Capgemini
|
Canada
Monitoring and Alerting Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users.
Respond to incidents and outages, diagnose problems, and implement solutions to minimize downtime and restore service.
Automation Automate repetitive tasks and processes to improve efficiency and reduce manual effort.
Infrastructure Management Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources.
Capacity Planning Plan for future capacity needs to ensure systems can handle anticipated workloads.
Release Engineering Develop and maintain processes for deploying software updates and releases.
Work closely with developers, operations teams, and other stakeholders to ensure system reliability and availability.
Documentation Maintain explicit and concise documentation of systems, processes, and procedures.
Identify areas for improvement and implement changes to enhance system reliability and performance.
Skills and Qualifications 8+ Years experience in production support handling Prod incidents.
Excellent knowledge of OCP and Azure Services.
Container Services (Kubernetes) is mandatory especially on Azure using CLI
Monitoring tools (Dynatrace), Splunk
Catchpoint
Knowledge of Chaos Engineering.
Scripting (Shell Scripting, Python, Power Shell)
Database (SQL database management, Mongo Db)
J-18808-Ljbffr
📌 Prod Support Sre Ontario (Canada)
🏢 Capgemini
📍 Canada