Elevate your career with a Cluster Site Reliability Engineer position, focusing on high-performance compute clusters and InfiniBand at our Cologix regions. This full-time role empowers you to own hardware operations and optimize cluster reliability. You will play a vital role in managing physical grill clusters while collaborating with a talented team of engineers across multiple locations.
Your responsibilities encompass bringing recent GPU racks online, driving InfiniBand fabric performance, and leading incident responses. With a focus on hands-on technology and problem resolution, you'll ensure the service level agreements are met, fostering system reliability in a rapid-paced environment. Key Responsibilities:
- Bring up new GPU racks with configurations and validations
- Drive InfiniBand fabric to achieve performance specs
- Conduct capacity planning across regions, forecast demand
- Own incident response, ensuring quick resolution targets
- Build and maintain operational runbooks for team efficiency Requirements:
- Over five years operating large compute clusters
- Expertise in InfiniBand and Ethernet RoCE networking
- Proficient in Go or Python scripting
- Ability to stay calm under pressure when resolving incidents
- Familiarity with bare-metal provisioning systems desirable Harness your expertise in compute clusters and network optimization with this exciting Cluster Site Reliability Engineer position.
📌 Cluster Site Reliability Engineer Role (Canada)
🏢 iFrame
📍 Canada