09 Aug
|
iFrame
|
Winnipeg
Elevate your career with a Cluster Site Reliability Engineer position, focusing on high-performance compute clusters and InfiniBand at our Cologix regions. This full-time role empowers you to own hardware operations and optimize cluster reliability.
You will play a vital role in managing physical grill clusters while collaborating with a talented team of engineers across multiple locations. Your responsibilities encompass bringing new GPU racks online, driving InfiniBand fabric performance, and leading incident responses. With a focus on hands-on technology and problem resolution, you'll ensure the service level agreements are met, fostering system reliability in a fast-paced environment.
Key Responsibilities:
• Bring up recent GPU racks with configurations and validations • Drive InfiniBand fabric to achieve performance specs • Conduct capacity planning across regions, forecast demand • Own incident response, ensuring quick resolution targets • Build and maintain operational runbooks for team efficiency
Requirements: • Over five years operating large compute clusters • Expertise in InfiniBand and Ethernet RoCE networking • Proficient in Go or Python scripting • Ability to stay calm under pressure when resolving incidents • Familiarity with bare-metal provisioning systems desirable
Harness your expertise in compute clusters and network optimization with this exciting Cluster Site Reliability Engineer position.
📌 Cluster Site Reliability Engineer Role (Winnipeg)
🏢 iFrame
📍 Winnipeg