Explore your skills in a Site Reliability Engineer position, focusing on the reliability of advanced compute clusters at our Cologix regions. This full-time role combines hands-on work with leadership in hardware operations and incident management.You will automatically engage with top-tier technology as part of a collaborative team overseeing multiple regions. Your role demands that you set up GPU racks, validate critical networking components, and maintain SLA standards through effective incident response management.
The emphasis on hands-on involvement means you will directly impact the efficiency and reliability of our computing services.Key Responsibilities:Build and test new GPU rack configurationsOptimize InfiniBand fabric for peak performancePerform capacity planning across regions while forecastingLead incident response efforts to ensure rapid resolutionCreate essential runbooks for operational proceduresRequirements:Minimum five years of large cluster operation experienceIn-depth knowledge of InfiniBand and networkingCompetency in Go or Python programmingResilience under pressure during incident managementNice to have: experience with provisioning systemsDrive innovation and assurance in compute cluster performance with this role as a Engineer in Site Reliability.#J-18808-Ljbffr
📌 Engineer In Site Reliability For Clusters (Ottawa)
🏢 iFrame
📍 Ottawa
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.