25 Sep
|
VySystems
|
Canada
Responsibilities
- Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You're responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth.
- Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, GCP, and OCI, with additional providers to be expanded in the near future.
- Manage job scheduling to allocate GPU compute across training and inference workloads.
- Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews.
- Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.
Requirements
- 4+ years in infrastructure engineering, cloud platforms,
or HPC.
- Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.
- Terraform proficiency. You'll write and review infrastructure-as-code daily.
- Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).
- Python for tooling and automation.
📌 Cloud Engineer (Canada)
🏢 VySystems
📍 Canada