Site Reliability Engineer, AI/ML Infrastructure (Toronto)

Site Reliability Engineer, AI/ML Infrastructure (Toronto)

06 Sep
|
Boson AI
|
Toronto

06 Sep

Boson AI

Toronto

re looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters aroundour Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers. Youll be hands-on with the full lifecycle of HPC infrastructure: planning, building, testing, deploying, and keeping everything running smoothly. That means troubleshooting issues as they arise, monitoring performance, developing automation to make our lives easier, and working closely with engineering and science teams to ensure they have what they need.

Youll also help us plan for future capacity and evaluate current technologies as we continue to scale.



Support ML/research teams with cluster usage optimization Proficiency in Linux systems administration (Ubuntu/Debian) ~ 1PB deployments and maintenance ~ Understanding of L2/L3 networking fundamentals ~ Skilled in Python and Bash scripting AWS, Azure or GCP We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment.

If you would like more information about how your data is processed, please contact us. #

📌 Site Reliability Engineer, AI/ML Infrastructure (Toronto)
🏢 Boson AI
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer, ai/ml infrastructure (toronto) / toronto