Elevate your career with a Cluster Site Reliability Engineer position, focusing on high-performance compute clusters and InfiniBand at our Cologix regions. This full-time role empowers you to own hardware operations and optimize cluster reliability.You will play a vital role in managing physical grill clusters while collaborating with a talented team of engineers across multiple locations. Your responsibilities encompass bringing new GPU racks online, driving InfiniBand fabric performance, and leading incident responses. With a focus on hands-on technology and problem resolution, you'll ensure the service level agreements are met,
fostering system reliability in a fast-paced environment.Key Responsibilities:Bring up current GPU racks with configurations and validationsDrive InfiniBand fabric to achieve performance specsConduct capacity planning across regions, forecast demandOwn incident response, ensuring quick resolution targetsBuild and maintain operational runbooks for team efficiencyRequirements:Over five years operating large compute clustersExpertise in InfiniBand and Ethernet RoCE networkingProficient in Go or Python scriptingAbility to stay calm under pressure when resolving incidentsFamiliarity with bare-metal provisioning systems desirableHarness your expertise in compute clusters and network optimization with this exciting Cluster Site Reliability Engineer position.#J-18808-Ljbffr
📌 Cluster Site Reliability Engineer Role (Ottawa)
🏢 iFrame
📍 Ottawa