Join a cutting-edge team as a Site Reliability Engineer, specializing in managing robust compute clusters and Infini Band infrastructure across Cologix regions. This full time, hands-on position is ideal for experienced engineers passionate about hardware and reliability.
In this role, you will take primary responsibility for your region's compute clusters and engage in cross-team collaboration. You will be involved in setting up and maintaining GPU racks, ensuring optimal Infini Band fabric performance to meet specifications, and planning for regional capacity needs.
This is more than operational oversight; it is a chance to significantly impact the reliability and efficiency of mission-critical systems.
Key Responsibilities:
Setup and validate GPU racks and configurations
Ensure Infini Band fabric performance is within required specs
Run regional capacity planning and demand forecasting
Manage incident responses, aiming for rapid resolution
Develop runbooks to enhance team knowledge transfer
Requirements:
Five-plus years managing large compute clusters
Robust background in Infini Band and Ethernet networking
Experience with Go or Python for scripting tasks
Ability to maintain composure under challenging situations
Additional familiarity with provisioning systems preferred
Take your skills in cluster management to the next level in this role focused on reliability and performance.
J-18808-Ljbffr
📌 Site Reliability Engineer For Compute Clusters Winnipeg (Canada)
🏢 iframe
📍 Canada
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.