28 Sep
|
eTeam
|
Mississauga
Role name: Dev
Ops Engineer (ML Infrastructure) Work Location: Canada (Remote) Dev
Ops Engineer (ML Infrastructure): We are seeking a highly skilled Senior Dev
Ops Engineer to help build, operate, and evolve large-scale, business-critical infrastructure platforms.
This role focuses on reliability, scalability, automation, and operational excellence across distributed systems supporting high-volume production workloads.
You will work closely with software engineers, platform teams, and infrastructure specialists to deliver highly available services, improve developer productivity, modernize infrastructure, and drive innovation through automation and AI-assisted operations.
This is an opportunity to solve complex technical challenges at scale while influencing the future direction of platform engineering and infrastructure management.
To design, build, and operate scalable machine learning infrastructure that enables productive training, deployment, and monitoring of AI/ML workloads.
The ideal candidate combines deep expertise in cloud-native technologies, Kubernetes, distributed systems, and software engineering to support large-scale machine learning platforms and GPU-based environments.
Key Responsibilities: Design, build, and maintain scalable MLOps platforms for training, deploying, and monitoring machine learning models.
Develop and manage cloud-native infrastructure supporting large-scale ML workloads on Kubernetes.
Implement and operate batch scheduling solutions such as Volcano and Kueue to optimize utilization of GPU and compute resources.
Build and maintain CI/CD pipelines for ML services, infrastructure, and platform components.
Automate infrastructure provisioning and lifecycle management using Infrastructure-as-Code practices.
Manage and optimize Kubernetes environments, including production workloads running on AWS EKS and other cloud platforms.
Lead and support infrastructure modernization and migration initiatives across cloud and platform ecosystems.
Partner with Data Scientists, ML Engineers, and Software Engineers to productionize machine learning solutions.
Implement robust observability, monitoring, and alerting for distributed systems and GPU clusters.
Ensure platform reliability, security, scalability, and operational excellence.
Troubleshoot complex distributed systems and performance bottlenecks across infrastructure and ML workloads.
Required Qualifications: years of experience in Dev
Ops, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering. years of experience supporting production machine learning or AI platforms.
Strong experience with Kubernetes and large-scale workload orchestration.
Hands-on experience with Kubernetes batch schedulers such as Volcano / Kueue Solid understanding of: Distributed systems Containerization technologies Cloud-native architectures Microservices-based platforms Experience managing workloads on Kubernetes platforms such as AWS EKS.
Proven track record delivering and supporting infrastructure migration projects.
Strong programming skills in Python or Golang Experience with Infrastructure-as-Code and deployment tools including Terraform / Helm Experience designing and maintaining CI/CD pipelines using Git
Hub Actions, Jenkins, Git
Lab CI/CD, or Azure Dev
Ops.
Strong Linux systems administration and troubleshooting skills.
Experience building observability solutions using: Prometheus Grafana Cloud-native monitoring tools Understanding of ML lifecycle management, model deployment, and production operations.
Preferred Qualifications: Experience supporting GPU-intensive machine learning or AI training platforms.
Hands-on experience with MLflow / Kubeflow Experience with distributed training frameworks and GPU resource management.
Familiarity with LLMOps, Generative AI, RAG architectures, and vector databases is a plus Knowledge of Git
Ops tools such as ArgoCD or Flux.
Experience with multi-cluster Kubernetes environments.
Parquet Mandatory skills: Distributed systems Containerization technologies Cloud-native architectures Microservices-based platforms CI/CD/Gitubs/Jenkkins,AWS EKS ML workflow Kubflow
📌 DevOps Engineer (ML Infrastructure) (Mississauga)
🏢 eTeam
📍 Mississauga