Elevate enterprise application performance as a Site Reliability Engineer with Hitachi. Your expertise in Kubernetes, cloud platforms, and operational excellence will be crucial in ensuring robust application stability. Join Hitachi's SRE Operations team to reinforce application reliability across cloud-native and hybrid platforms. This position requires 2-5 years of experience in IT Operations or SRE, a solid understanding of Kubernetes, and proficiency in Linux troubleshooting.
Key responsibilities include monitoring infrastructure, performing incident triage, and collaborating with engineering teams to ensure service reliability. Key Responsibilities:
- Monitor applications and infrastructure across environments
- Perform incident triage and execute operational runbooks
- Troubleshoot application issues using Linux utilities
- Support Kubernetes deployment and validation of pod health
- Maintain incident documentation and knowledge base updates Requirements:
- 2–5 years in IT Operations, SRE, or DevOps
- Solid Linux administration knowledge
- Familiarity with AWS, Azure, or GCP
- Experience with monitoring tools like Prometheus
- Basic scripting skills in Python or Bash Bring your passion for automation and operational excellence to Hitachi's innovative SRE team.
📌 Site Reliability Engineer at Hitachi (Toronto)
🏢 Socket.dev
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.