Principal AI Cloud Engineer (Toronto)

Principal AI Cloud Engineer (Toronto)

02 Oct
|
BMO Financial
|
Toronto

02 Oct

BMO Financial

Toronto

Data Analytics & Reporting

The Team - We accelerate BMO’s AI journey by building cloud-native AI solutions. Our team combines engineering excellence with cutting-edge AI to deliver scalable, secure, and responsible solutions that power business innovation across the bank. We enable and accelerate our partners on their AI journeys across the enterprise, helping teams across BMO unlock value at scale. We support one another in times of need and take pride in our work. We are engineers, AI practitioners, platform builders, thought leaders, multipliers, and coders. Our ambition is bold: deploy our capital and resources to their highest and most profitable use through a digital-first operating model, powered by data and AI-driven decisions.

As a Principal AI & Cloud Engineer , you are a hands-on technical developer who designs, builds, and scales cloud-native AI solutions and products. You help set engineering standards, establish patterns, mentor senior engineers, and partner with multiple teams to deliver resilient, governed, and cost-efficient AI at enterprise scale. You’ll help shape and evolve our AI cloud strategy from model serving and LLMOps to security, observability, and compliance so teams across the bank can innovate safely and rapidly.

You will advance BMO’s Digital First strategy by:
Defining reference and production-grade solutions for AI/GenAI on cloud (Azure/AWS preferred; Building reusable, secure, and observable components (APIs, SDKs, microservices, pipelines).
Operationalizing LLMs and RAG with strong controls and Responsible AI guardrails.
Driving platform roadmaps that enable faster delivery, lower risk, and measurable business outcomes.

Influence the technical direction of AI and the platform primitives others build on.
Ship high-impact systems used across many business lines and products.
Work across the full stack: cloud infra, data/feature pipelines, model serving, LLMOps, and DevSecOps.
Design, build,



and operate cloud-native AI infrastructure for ML/GenAI workloads:
Networking: Azure VNet, Private Link, peering, multi-region HA/DR
Storage & Databases: high-performance data lakes (e.g., Azure Data Lake Storage) , relational DBs, vector DBs (FAISS, Milvus, Pinecone, pgvector)
Security: IAM, Key Vault-backed secrets management, encryption, policy-as-code
Implement observability and reliability for AI infra:
Build CI/CD and GitOps pipelines for infrastructure-as-code (Terraform/Bicep) and AI platform components
Drive FinOps for AI infra: Enable frontend and backend services for AI platforms:
Provide infrastructure support for RAG systems: embeddings, chunking, retrieval pipelines
Ensure scalable serving infrastructure for LLMs and ML models with caching and token optimization
Define and evolve AI infrastructure reference architecture for cloud (Azure preferred):
Serverless/event-driven patterns for AI pipelines
Establish standards and best practices for containerization, IaC, and secure networking for AI systems
Security, Risk & Governance
Implement defense-in-depth for AI infra:
IAM least privilege, private networking, KMS/Key Vault, SBOM, image signing
Ensure compliance and Responsible AI controls at infra level:
Data residency, encryption, lineage, audit readiness
Operate platforms with SRE principles: error budgets, incident response, chaos testing
Bachelor’s/Master’s/PhD in CS, Engineering, or related field
~7+ years building large-scale distributed cloud infrastructure
~5+ years hands-on with Azure/AWS
~ Proven experience with AI/ML infra: Strong in IaC (Terraform/Bicep), Kubernetes,



networking, security
~ Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs
~ Programming in Python (infra automation) and one of Go/TypeScript for tooling
~ Understanding of frontend/backend integration for AI services
~ Familiarity with MLOps/LLMOps infra: model serving, feature stores, vector DBs
~ Programming in Python (infra automation) and one of Go/TypeScript for tooling
~ Understanding of frontend/backend integration for AI services

Event streaming (Kafka/Azure Event Hubs), real-time systems
Experience with AI platform products (Azure ML, MLflow, KServe, Hugging Face)
Reliability & Performance: SLOs met for infra services, GPU utilization optimized
Developer Velocity: Faster provisioning and deployment of AI infra
Technical Leadership: Salaries for part-time roles will be pro-rated based on number of hours regularly worked. BMO Financial Group’s total compensation package will vary based on the pay type of the position and may include performance-based incentives, discretionary bonuses, as well as other perks and rewards. BMO also offers health insurance, tuition reimbursement, accident and life insurance, and retirement savings plans. It calls on us to create lasting, positive change for our customers, our communities and our people. We strive to help you make an impact from day one – for yourself and our customers. We’ll support you with the tools and resources you need to reach current milestones, as you help our customers reach theirs. From in-depth training and coaching, to manager support and network-building opportunities, we’ll help you gain valuable experience, and broaden your skillset.

Accommodations are available on request for candidates taking part in all aspects of the selection process. A recruiting agency must first have a valid, written and fully executed agency agreement contract for service to submit resumes.

📌 Principal AI Cloud Engineer (Toronto)
🏢 BMO Financial
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal ai cloud engineer (toronto) / toronto