22 Aug
|
iframe
|
Winnipeg
Join a cutting-edge team as a Site Reliability Engineer, specializing in managing robust compute clusters and Infini Band infrastructure across Cologix regions. This full time, hands-on position is ideal for experienced engineers passionate about hardware and reliability.
In this role, you will take primary responsibility for your region's compute clusters and engage in cross-team collaboration. You will be involved in setting up and maintaining GPU racks, ensuring optimal Infini Band fabric performance to meet specifications, and planning for regional capacity needs. This is more than operational oversight;
it is a chance to significantly impact the reliability and efficiency of mission-critical systems.
Key Responsibilities:
Setup and validate GPU racks and configurations
Ensure Infini Band fabric performance is within required specs
Run regional capacity planning and demand forecasting
Manage incident responses, aiming for rapid resolution
Develop runbooks to enhance team knowledge
📌 Site Reliability Engineer For Compute Clusters Winnipeg
🏢 iframe
📍 Winnipeg