Senior SRE: AI/ML HPC Infra & GPU Cluster (Ontario)

Senior SRE: AI/ML HPC Infra & GPU Cluster (Ontario)

10 Aug
|
Boson AI
|
Ontario

10 Aug

Boson AI

Ontario

A technology company in Toronto seeks a Senior Site Reliability Engineer to manage and optimize its HPC infrastructure. In this role, you'll ensure smooth operations of a powerful GPU cluster, deploy infrastructure-as-code solutions, and support ML teams. Candidates should have extensive SRE experience, proficiency in Linux, and familiarity with Kubernetes and Ceph storage. This position offers the chance to work with cutting-edge technology in a team-oriented environment, perfect for problem-solvers who love learning.
#J-18808-Ljbffr

📌 Senior SRE: AI/ML HPC Infra & GPU Cluster (Ontario)
🏢 Boson AI
📍 Ontario

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior sre: ai/ml hpc infra & gpu cluster (ontario) / ontario

Subscribe to this job alert:

Get the latest job offers by email for: senior sre: ai/ml hpc infra & gpu cluster (ontario) / ontario