Elevate your career with a Cluster Site Reliability Engineer position, focusing on high-performance compute clusters and InfiniBand at our Cologix regions. This full time role empowers you to own hardware operations and optimize cluster reliability.
You will play a vital role in managing physical grill clusters while collaborating with a talented team of engineers across multiple locations. Your responsibilities encompass bringing new GPU racks online, driving InfiniBand fabric performance, and leading incident responses. With a focus on hands-on technology and problem resolution, you'll ensure the service level agreements are met, fostering system reliability in a fast-paced environment.
Key Responsibilities:
• Bring up recent GPU racks with configurations and validations
• Drive InfiniBand fabric to achieve performance specs
• Conduct capacity planning across regions, forecast demand
• Own incident response, ensuring quick resolution targets
• Build and maintain operational runbooks for team efficiency
Requirements:
• Over five years operating large compute clusters
• Expertise in InfiniBand and Ethernet RoCE networking
• Proficient in Go or Python scripting
• Ability to stay calm under pressure when resolving incidents
• Familiarity with bare-metal provisioning systems desirable
Harness your expertise in compute clusters and network optimization with this exciting Cluster Site Reliability Engineer position.
#J-18808-Ljbffr
📌 Cluster Site Reliability Engineer Role (Manitoba)
🏢 iFrame
📍 Manitoba
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.