Explore your skills in a Site Reliability Engineer position, focusing on the reliability of advanced compute clusters at our Cologix regions. This full time role combines hands-on work with leadership in hardware operations and incident management.
You will automatically engage with top-tier technology as part of a collaborative team overseeing multiple regions. Your role demands that you set up GPU racks, validate critical networking components, and maintain SLA standards through effective incident response management. The emphasis on hands-on involvement means you will directly impact the efficiency and reliability of our computing services.
Key Responsibilities:
• Build and test new GPU rack configurations
• Optimize InfiniBand fabric for peak performance
• Perform capacity planning across regions while forecasting
• Lead incident response efforts to ensure fast resolution
• Create essential runbooks for operational procedures
Requirements:
• Minimum five years of large cluster operation experience
• In-depth knowledge of InfiniBand and networking
• Competency in Go or Python programming
• Resilience under pressure during incident management
• Nice to have: experience with provisioning systems
Drive innovation and assurance in compute cluster performance with this role as a Engineer in Site Reliability.
#J-18808-Ljbffr
📌 Engineer in Site Reliability for Clusters (Manitoba)
🏢 iFrame
📍 Manitoba
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.