17 Aug
|
Avenue Code
|
Toronto
17 Aug
Avenue Code
Toronto
About the Opportunity
Avenue Code is seeking a Senior AI/ML Infrastructure Engineer to support the development and evolution of a large-scale Machine Learning Platform for a global technology company.
This role will focus on building the infrastructure, tooling, and platform capabilities that enable Machine Learning Engineers, researchers, and product teams to efficiently train, deploy, and serve AI and ML models at scale.
You will work on production-grade ML infrastructure, contributing to platform SDKs and developer tooling while managing Kubernetes-based environments designed to support demanding machine learning workloads, including distributed training and GPU-based compute.
This is a highly technical, infrastructure-focused opportunity ideal for engineers with strong software engineering foundations who enjoy working at the intersection of Machine Learning, Platform Engineering, Kubernetes, and distributed systems.
This is a hybrid role based in Toronto, Canada. Candidates must be currently based in Canada and able to attend occasional in-person meetings in Toronto.
Responsibilities
- Design, build, and maintain scalable infrastructure and tooling for training and serving Machine Learning models in production.
- Contribute to ML Platform SDKs and develop tools supporting various ML operations and workflows.
- Build reliable and scalable platform capabilities that simplify the productionization of AI and ML models.
- Manage and maintain large-scale production Kubernetes clusters supporting ML workloads.
- Support infrastructure for distributed model training, including GPU-based workloads.
- Collaborate closely with Machine Learning Engineers, researchers,
and product engineering teams to deliver scalable ML platform solutions.
- Design, document, implement, and maintain reliable, testable, and maintainable ML infrastructure capabilities.
- Support the operational reliability, scalability, and performance of ML platform infrastructure.
- Evaluate and apply new technologies to solve evolving infrastructure and ML platform challenges.
- Contribute to engineering best practices around modular architecture, testing, automation, and production operations.
Required Qualifications
- 6+ years of hands‑on software engineering, ML Infrastructure, Platform Engineering, or related experience, with demonstrated experience implementing production ML infrastructure at scale.
- Strong programming experience with Python, Go, or similar backend languages.
- Hands‑on experience managing Kubernetes environments in production.
- Experience supporting Machine Learning workloads and infrastructure at scale.
- Understanding of distributed model training, including GPU‑based workloads and Kubernetes‑based execution.
- Knowledge of deep learning fundamentals, algorithms, and up-to-date ML ecosystems.
- Experience with ML frameworks and tools such as PyTorch, TensorFlow, Hugging Face, Ray, or similar technologies.
- General understanding of data processing and data workflows for Machine Learning.
- Strong understanding of software engineering principles, modular code design, testing, and maintainability.
- Experience working within Agile software development environments.
- Strong communication and collaboration skills, with the ability to work effectively with Machine Learning Engineers, researchers, and product teams.
Nice To Have Skills
- Experience building or maintaining internal Machine Learning Platforms or ML developer platforms.
- Experience developing SDKs, APIs, or developer tooling for Machine Learning workflows.
- Deep experience with GPU infrastructure and large‑scale distributed training.
- Experience with ML model serving and production inference infrastructure.
- Familiarity with MLOps practices and model lifecycle management.
- Experience with cloud‑native infrastructure supporting large‑scale AI/ML workloads.
- Experience improving developer productivity and simplifying ML productionization workflows.
A reasonable estimate of the current range for this position is CAD $165k to CAD $175k yearly.
Avenue Code reinforces its commitment to privacy and to all the principles guaranteed by the most accurate global data protection laws, such as GDPR, LGPD, CCPA and CPRA. The Candidate data shared with Avenue Code will be kept confidential and will not be transmitted to disinterested third parties, nor will it be used for purposes other than the application for open positions. As a Consultancy company, Avenue Code may share your information with its clients and other Companies from the CompassUOL Group to which Avenue Code’s consultants are allocated to perform its services.
📌 AI/ML Infrastructure Engineer (Toronto)
🏢 Avenue Code
📍 Toronto