25 Sep
|
Umanist Staffing
|
Canada
25 Sep
Umanist Staffing
Canada
HPC Consultant – Kubernetes / GPU / AWS
Job Type: Full time
Location: Remote
Work Model: Remote-first; occasional on-site visits to customer offices
Job Summary
We are looking for an experienced HPC Consultant / Cloud Infrastructure Engineer with strong hands-on expertise in Kubernetes, HPC/GPU infrastructure, Terraform, and AWS.
This is a Kubernetes-heavy infrastructure role supporting large-scale, multi-cloud GPU platforms. The ideal candidate will have experience operating production Kubernetes clusters at meaningful scale, troubleshooting complex cluster issues, automating infrastructure through Infrastructure-as-Code, and supporting high-performance computing workloads.
Key Responsibilities
- Operate production Kubernetes platforms including EKS, CKS, and GKE.
- Manage Kubernetes cluster lifecycle, node pools, upgrades, networking policies, and platform stability.
- Troubleshoot complex Kubernetes issues including:
- Scheduler problems
- CNI/networking issues
- Node failures
- Resource allocation
- Cluster upgrades
- Provision and manage HPC/GPU infrastructure using CI/CD and Terraform.
- Support infrastructure across AWS, GCP, OCI,
CoreWeave, and other cloud providers.
- Manage GPU compute resources for AI/ML training and inference workloads.
- Monitor cluster and infrastructure health and maintain SLIs/SLOs.
- Build and maintain monitoring, alerting, and operational dashboards.
- Participate in production incident response and severity escalations.
- Conduct post-incident reviews and implement corrective actions.
- Work closely with Networking, Storage, Security, and AI/ML Platform teams.
- Develop Python-based tooling and automation for infrastructure operations.
Required SkillsKubernetes – MUST HAVE
- 4+ years of infrastructure engineering, cloud platform, or HPC experience.
- Strong hands-on production Kubernetes experience.
- Experience operating large-scale Kubernetes clusters.
- Experience with
- EKS / GKE / CKS
- Node pool management
- Kubernetes scheduling
- CNI/networking troubleshooting
- Rolling upgrades
- Cluster lifecycle management
- Production troubleshooting
📌 HPC Consultant – Kubernetes / GPU / AWS (Canada)
🏢 Umanist Staffing
📍 Canada