- Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
- Automate workflows using Go, Python, and Shell scripting.
- Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
- Troubleshoot complex networking, storage, and system performance issues.
- Partner with AI/ML teams to ensure infrastructure readiness for model training and data pipelines.
- Participate in on-call rotations and postmortem reviews to improve system resilience.
Qualifications
- Experience with Google Cloud, plus IaC tools (Terraform).
- Solid knowledge of microservices, containers (Kubernetes, Docker), and networking.
- Hands‑on experience with PKI, service mesh, and Linux systems administration.
- SRE mindset with a focus on automation, scalability, and reliability.
- Experience with Golang is an asset.
Benefits
- Competitive total rewards package, with paid vacation, sick days, and a day off to volunteer.
- Training allowance and professional development opportunities.
- Home‑office equipment and an annual budget to personalize your work environment.
- Annual wellness budget for gym memberships, massages, and fitness.
Pay Range
90,000 - 100,000 CAD per year (Canada)
Hiring Disclaimer
- The successful applicant will need to fulfill the requirements necessary to obtain a background check.
- Accommodations are available upon request for candidates taking part in any aspect of the selection process.
- Position is active and open; we are looking to fill it as soon as possible.
#J-18808-Ljbffr
📌 Site Reliability Consultant (Vancouver)
🏢 Pythian
📍 Vancouver
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.