09 Oct
|
EPAM Systems
|
Canada
09 Oct
EPAM Systems
Canada
Join a high-growth infrastructure team operating Kubernetes platforms across multiple cloud providers at massive scale.
You'll build the systems that power thousands of GPUs, where your code and configurations directly protect thousands of GPU-hours from costly failures.
EPAM is where tech talent thrives—building groundbreaking solutions, advancing your skills through world-class learning platforms, and working alongside a global community of problem-solvers to make the future real.
Req# (phone hidden)
**Responsibilities**
Operate and scale Kubernetes platforms (EKS, GKE, and other distributions) including cluster lifecycle management, node pool optimization, and networking policies during periods of rapid growth
Provision and manage HPC infrastructure through CI/CD pipelines spanning AWS, Core
Weave, GCP, OCI, and additional cloud providers
Design and maintain job scheduling systems that efficiently allocate GPU compute resources across training and inference workloads
Define SLIs/SLOs, build robust monitoring and alerting systems, and actively participate in incident response and post-incident reviews
Develop production-quality tooling and automation to support multi-cloud infrastructure operations at scale
Collaborate daily with Networking, Storage, Security, and AI/ML platform teams to ensure seamless cross-functional infrastructure delivery
**Requirements**
10+ years of experience in infrastructure engineering, cloud platforms, or high-performance computing environments
Expert-level Kubernetes experience at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
Advanced Python skills with a track record of building production-grade tools, not just scripts; experience with Go, Rust, or C++ is a solid plus
Daily proficiency in Terraform for writing and reviewing infrastructure as code
Working knowledge of core AWS services including EC2, S3, EFS, and FSx for Lustre
Strong site reliability engineering background with experience building monitoring, alerting, and incident response practices
CKA, CKS certificates are highly preferred
**We offer**
Extended Healthcare with Prescription Drugs, Dental and Vision, and Healthcare Spending Account (Company Paid)
Life and AD&D; Insurance (Company Paid)
Employee Assistance Program (Company Paid)
Telehealth (Company Paid)
Short-term Disability (Company Paid)
Long-Term Disability
Paid Time Off (including vacation and sick days)
Registered Retirement Savings Plan (RRSP) with Company match
Maternity/Parental/Adoption Leave Top-up
Employee Stock Purchase Program
Critical Illness Insurance
Employee Discounts
Unlimited access to Linked
In learning solutions
EPAM Systems, Inc. is an equal opportunity employer.
We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities.
We are committed to recruiting, hiring, developing and promoting employees without discrimination.
As a global employer, this commitment includes complying with all laws in the countries in which we operate.
Nevertheless, we believe equal employment practices should not be limited to what the law requires.
Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.
At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.
📌 Lead Platform Engineer/Architect - HPC, Kubernetes (Canada)
🏢 EPAM Systems
📍 Canada