Engineer In Site Reliability For Clusters Winnipeg

Engineer In Site Reliability For Clusters Winnipeg

25 Aug
|
iframe
|
Winnipeg

25 Aug

iframe

Winnipeg

Explore your skills in a Site Reliability Engineer position, focusing on the reliability of advanced compute clusters at our Cologix regions. This full-time role combines hands-on work with leadership in hardware operations and incident management. You will automatically engage with top-tier technology as part of a cooperative team overseeing multiple regions.

Your role demands that you set up GPU racks, validate critical networking components, and maintain SLA standards through effective incident response management. The emphasis on hands-on involvement means you will directly impact the efficiency and reliability of our computing services. Key Responsibilities:
Build and test current GPU rack configurations
Optimize InfiniBand fabric for peak performance




Perform capacity planning across regions while forecasting
Lead incident response efforts to ensure quick resolution
Create essential runbooks for operational procedures Requirements:
Minimum five years of large cluster operation experience
In-depth knowledge of InfiniBand and networking
Competency in Go or Python programming
Resilience under pressure during incident management
Nice to have: experience with provisioning systems Drive innovation and assurance in compute cluster performance with this role as a Engineer in Site Reliability.

📌 Engineer In Site Reliability For Clusters Winnipeg
🏢 iframe
📍 Winnipeg

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: engineer in site reliability for clusters winnipeg / winnipeg

Subscribe to this job alert:

Get the latest job offers by email for: engineer in site reliability for clusters winnipeg / winnipeg