21 Sep
|
AceStack
|
Toronto
Job Title: Platform Engineer - AI/ML Infrastructure Job Type: Full Time Location: Toronto, ON Work Model: 4 Days/Week Primary Skills: Kubernetes, Git
Ops (Flux CD/ArgoCD), Helm & Hashi
Corp Vault Job Description We are seeking a skilled Platform Engineer - AI/ML Infrastructure to deploy, manage, and support reliable, scalable, and secure AI/ML infrastructure across multiple environments.
The ideal candidate will have strong hands-on experience with Kubernetes, Git
Ops practices, CI/CD pipelines, secrets management, and cloud-native infrastructure.
The successful candidate will work with AI/ML platforms such as Llama
Index Cloud and KDB.AI, supporting deployments across Development, QA, and Production environments while ensuring platform reliability, security, and operational efficiency.
Roles & Responsibilities Deploy and manage Llama
Index Cloud, KDB.AI, and similar AI/ML applications across Dev, QA, and Production environments.
Build reliable, scalable, and secure deployment pipelines using up-to-date Git
Ops practices.
Implement and maintain Git
Ops workflows using Flux CD for automated deployments.
Administer and maintain Kubernetes clusters across multiple environments.
Configure Hashi
Corp Vault for secrets management and integrate with External
Secrets.
Maintain CI/CD pipelines using Git
Hub Actions, Jenkins, or similar tools.
Work with the enterprise Identity Management team to configure OIDC authentication with Microsoft Entra ID.
Create, maintain, and optimize Helm charts and Kubernetes manifests.
Configure and manage Kubernetes networking, Ingress, and Gateway API.
Monitor application performance and troubleshoot production issues.
Support PostgreSQL, MongoDB, Redis,
and RabbitMQ infrastructure as required.
Implement and maintain database high-availability and failover configurations.
Develop and maintain infrastructure documentation, runbooks, and operational procedures.
Collaborate with engineering, security, identity, and platform teams to improve infrastructure reliability and automation.
Required Skills & Qualifications 3+ years of experience managing production Kubernetes clusters. 2+ years of experience with Flux CD, ArgoCD, or similar Git
Ops tools .
Advanced experience developing and managing Helm charts.
Strong hands-on experience with Hashi
Corp Vault for secrets management.
Experience with Artifactory or similar container registries.
Strong experience with CI/CD tools such as Git
Hub Actions or Jenkins.
Administration experience with PostgreSQL, MongoDB, Redis, and RabbitMQ.
Experience with database HA and failover configurations, including PgBouncer and HAProxy.
Strong Linux/Unix administration skills and shell scripting using Bash or Power
Shell.
Good understanding of Kubernetes networking, Ingress, and Gateway API.
Strong troubleshooting, monitoring, and problem-solving skills.
Nice-to-Have Skills Experience with Llama
Index, Lang
Chain, or other AI/ML platforms.
Knowledge of vector databases or KDB.AI.
Experience with Temporal.io workflow orchestration.
Programming experience with Python or Go for infrastructure automation.
Experience supporting production AI/ML workloads.
Key Competencies Kubernetes Administration Git
Ops & Continuous Deployment Flux CD / ArgoCD Helm Hashi
Corp Vault CI/CD & Dev
Ops AI/ML Infrastructure Linux/Unix Infrastructure Automation Production Support
📌 Platform Engineer - AI/ML Infrastructure (Toronto)
🏢 AceStack
📍 Toronto