Join a cutting-edge team as a Site Reliability Engineer, specializing in managing robust compute clusters and InfiniBand infrastructure across Cologix regions. This full-time, hands-on position is ideal for experienced engineers passionate about hardware and reliability.
In this role, you will take primary responsibility for your region's compute clusters and engage in cross-team collaboration. You will be involved in setting up and maintaining GPU racks, ensuring optimal InfiniBand fabric performance to meet specifications, and planning for regional capacity needs. This is more than operational oversight; it is a chance to significantly impact the reliability and efficiency of mission-critical systems.
Key Responsibilities:
• Setup and validate GPU racks and configurations
• Ensure InfiniBand fabric performance is within required specs
• Run regional capacity planning and demand forecasting
• Manage incident responses, aiming for rapid resolution
• Develop runbooks to enhance team knowledge transfer
Requirements:
• Five-plus years managing large compute clusters
• Solid background in InfiniBand and Ethernet networking
• Experience with Go or Python for scripting tasks
• Ability to maintain composure under challenging situations
• Additional familiarity with provisioning systems preferred
Take your skills in cluster management to the next level in this role focused on reliability and performance.
#J-18808-Ljbffr
📌 Site Reliability Engineer for Compute Clusters (Manitoba)
🏢 iFrame
📍 Manitoba
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.