Regional Cluster SRE — GPU/HPC Infra Lead (Ontario)

Regional Cluster SRE — GPU/HPC Infra Lead (Ontario)

13 Aug
|
iframe
|
Ontario

13 Aug

iframe

Ontario

iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This in office role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA while being the on-call owner for regional incidents.
You will work with Linux, Kubernetes, Terraform, Prometheus/Grafana, and Go or Python, coordinating with procurement and regional teams.

#J-18808-Ljbffr

📌 Regional Cluster SRE — GPU/HPC Infra Lead (Ontario)
🏢 iframe
📍 Ontario

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: regional cluster sre — gpu/hpc infra lead (ontario) / ontario

Subscribe to this job alert:

Get the latest job offers by email for: regional cluster sre — gpu/hpc infra lead (ontario) / ontario