Elevate your career with a Cluster Site Reliability Engineer position, focusing on high-performance compute clusters and Infini Band at our Cologix regions. This full-time role empowers you to own hardware operations and optimize cluster reliability.
You will play a vital role in managing physical grill clusters while collaborating with a talented team of engineers across multiple locations. Your responsibilities encompass bringing new GPU racks online, driving Infini Band fabric performance, and leading incident responses.
With a focus on hands-on technology and problem resolution, you'll ensure the service level agreements are met, fostering system reliability in a fast-paced environment.
Key Responsibilities:
- Bring up current GPU racks with configurations and validations
- Drive Infini Band fabric to achieve performance specs
- Conduct capacity planning across regions, forecast demand
- Own incident response, ensuring quick resolution targets
- Build and maintain operational runbooks for team efficiency
Requirements:
- Over five years operating large compute clusters
- Expertise in Infini Band and Ethernet RoCE networking
- Proficient in Go or Python scripting
- Ability to stay calm under pressure when resolving incidents
- Familiarity with bare-metal provisioning systems desirable
Harness your expertise in compute clusters and network optimization with this exciting Cluster Site Reliability Engineer position.
#J-18808-Ljbffr
📌 Cluster Site Reliability Engineer Role (Winnipeg)
🏢 iFrame
📍 Winnipeg
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.