14 Aug
|
iframe
|
Toronto
iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up current racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.
You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.
#J-18808-Ljbffr
📌 Regional Cluster Sre — Gpu/Hpc Infra Lead (Toronto)
🏢 iframe
📍 Toronto