- 8+ years of experience in production support handling production incidents.
- Excellent knowledge of OpenShift (OCP) and Windows environments.
- Proficiency with monitoring tools such as Dynatrace.
- Experience with Operating Systems including Windows and Linux.
- Proficiency in scripting languages such as Shell Scripting, Python, and PowerShell.
- Experience with database management including SQL and MongoDB.
- Experience with container services such as Kubernetes.
- Experience with Disaster Recovery planning and execution.
Responsibilities:
- Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users.
- Respond to incidents and outages, diagnose problems, and implement solutions to minimize downtime and restore service.
- Automate repetitive tasks and processes to improve efficiency and reduce manual effort.
- Manage and maintain the underlying infrastructure, including servers, networks,
and cloud resources.
- Plan for future capacity needs to ensure systems can handle anticipated workloads.
- Develop and maintain processes for deploying software updates and releases.
- Work closely with developers, operations teams, and other stakeholders to ensure system reliability and availability.
- Maintain transparent and concise documentation of systems, processes, and procedures.
- Identify areas for improvement and implement changes to enhance system reliability and performance.