13 Aug
|
iframe
|
Ontario
iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This in office role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.
You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.
#J-18808-Ljbffr
📌 Regional Cluster SRE — GPU/HPC Infra Lead (Ontario)
🏢 iframe
📍 Ontario