Enhance operational integrity as a Site Reliability Consultant with expertise in Kubernetes and automation. Focus on building and optimizing high-performance systems.
You will manage Kubernetes clusters, leverage Istio, and implement automation strategies using Go, Python, and Shell scripting. Responsibilities include diagnosing complex issues, constructing monitoring solutions with Prometheus and Grafana, and collaborating with AI/ML teams to prepare infrastructures. Engage in on-call rotations and participate in postmortem reviews to boost system resilience.
Key Responsibilities:
• Manage Kubernetes clusters and utilize Istio service mesh
• Create automated workflows with Go, Python, and Shell scripting
• Set up monitoring solutions with Prometheus, Grafana, and Loki
• Troubleshoot complex networking and storage performance issues
• Partner with AI/ML teams to ensure model readiness
Requirements:
• Experience with Google Cloud and Terraform tools
• Proficient in microservices and container technologies
• Robust Linux systems administration and PKI knowledge
• Cultivated SRE mindset geared towards scalability
• Golang experience considered a plus
Apply your expertise in Kubernetes and automation to significantly improve our system reliability.
J-18808-Ljbffr
📌 Consultant In Site Reliability And Automation British Columbia (Canada)
🏢 Pythian
📍 Canada
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.