AI Infrastructure & Solutions Architect (Toronto)

AI Infrastructure & Solutions Architect (Toronto)

03 Oct
|
University of Toronto
|
Toronto

03 Oct

University of Toronto

Toronto

Date Posted: 09/29/2026 Faculty/Division: VP-People, Finance &

• Digital Services The Enterprise Infrastructure Solutions (EIS) group, part of the Information Technology Services (ITS) division, is responsible for campus core network, campus wireless, wide area network connectivity and internet connectivity for the University, including connectivity to research and education networks. EIS is also responsible for services related to departmental network management, network, server and storage infrastructure, Windows and Linux servermanagement services, database and application integration and support, enterprise backup service, 24/7 operation of central administrative data centres and telecommunications services. If you’re motivated and passionate about learning technologies and dedicated to improving experiences for today’s student, consider a career with us. Reporting to the Manager, AI Engineering &

• Operations within the Enterprise Infrastructure Solutions group, the AI Infrastructure &

• Solutions Architect plays a critical role in defining the future of research and administrative computing at the University. In this role, you will lead the architectural design and deployment of secure, scalableAI platforms that serve the entire campus community, spanning the University’s own data centres and private cloud as well as the major public cloud platforms (AWS, Microsoft Azure, and Google Cloud). You will bridge the gap between high-performancehardware, cloud services, and practical user applications, ensuring that our AI platforms, from on-premises GPU clusters and sovereign data sandboxes to cloud-hosted AI services, are reliable, supportable, cost-effective, and aligned with institutional governance and ethical AI frameworks. You will collaborate closely with AI Developers &

• Integration Specialists, while retaining ownership of platform architecture, workload placement, reliability, security posture, cost efficiency, and lifecycle management across the University’s on-premises infrastructure and multiple public cloud environments. This role offers a rare opportunity to design and operate AI platforms at institutional scale across a multi-cloud estate that combines sovereign, on-premises capabilities with AWS, Microsoft Azure, and Google Cloud, supporting both cutting-edge research and mission-critical administrative use cases. EKS, AKS, GKE), including GPU operators and AI-aware schedulers for efficient accelerator utilization. Defining and maintaining the University’s multi-cloud AI reference architecture and workload placement framework, determining where AI workloads run across the University’s on-premises infrastructure, AWS, Microsoft Azure, and Google Cloud based on data classification, sovereignty, cost, performance, and service availability. Designing and integrating managed cloud AI services (e.g., Amazon Bedrock and SageMaker, Azure AI Foundry and Azure OpenAI, Google Vertex AI) with on-premises platforms through common abstractions such as model gateways, federated identity, and cloud landing zones provisioned through Infrastructure as Code. Designing and enforcing AI platform security and governance controls across on-premises and cloud environments,



including data sovereignty and residency (e.g., Canadian cloud regions), identity federation and access isolation, auditability, and compliance with privacy and ethical AI frameworks. Implementing observability and monitoring solutions to track model performance, drift, GPU utilization, inference latency, cost, and platform health using tools such as Prometheus, Grafana, OpenTelemetry, cloud-native monitoring services (e.g., Amazon CloudWatch, Azure Monitor, Google Cloud Monitoring), or specialized AI monitoring stacks, providing a unified view across on-premises and cloud platforms. NVIDIA, AMD, cloud GPU and TPU instances, or emerging accelerators), including capacity planning, scheduling strategies, burst-to-cloud patterns, and lifecycle management. Analyzing platform usage and cost metrics across on-premises and cloud platforms to optimize token consumption, GPU allocation, cloud spend (FinOps), and overall cost efficiency while maintaining performance and reliability. Partnering with AI Developers &

• Integration Specialists to define platform abstractions, deployment patterns, and service interfaces that enable rapid innovation without compromising security, supportability, or portability across on-premises and cloud environments. Producing and maintaining architectural documentation, disaster recovery and business continuity plans (including cross-cloud and cloud-to-on-premises recovery strategies), and technical guidance for researchers and platform users. Bachelor’s degree in Computer Science, Information Technology, Engineering, or an acceptable combination of education and equivalent experience. Eight or more years of experience in on-premises and public cloud infrastructure management. Three to five+ years of direct AI infrastructure or MLOps experience, recognizing the rapid evolution of the field, with demonstrated exposure to LLM platforms, RAG pipelines, or large-scale ML systems. MLOps and model lifecycle tooling experience, such as MLflow, Kubeflow, Weights &

• Biases, and model serving frameworks like Triton or vLLM, as well as their cloud-managed equivalents (e.g., Amazon SageMaker, Azure Machine Learning, Vertex AI). Strong understanding of high-performance and AI-optimized networking, including high-throughput, low-latency designs (e.g., AWS Direct Connect, Azure ExpressRoute,Google Cloud Interconnect, private endpoints, and cross-cloud networking). Demonstrated experience designing and operating hybrid and multi-cloud architectures, with hands-on expertise in at least two of the three major public cloud platforms (AWS, Microsoft Azure, Google Cloud) and their AI/ML and GPU compute services, including identity federation, networking, security, and cost management (FinOps) considerations. Assets (Nonessential):



Experience with vector databases and retrieval systems supporting RAG-based architectures. Hands-on experience deploying open-weights AI models (e.g., Knowledge of AI security practices, including adversarial ML considerations, secure AI framework implementations, or cloud security posture management (CSPM) across multiple providers. Familiarity with Canadian data sovereignty, privacy, and research compliance requirements, particularly within higher education, including the data residency options offered by Canadian cloud regions. Contributions to open-source AI, infrastructure, or MLOps projects, or active participation in the AI engineering community. ServiceNow) and operational workflow automation. AWS Certified Solutions Architect – Professional, Microsoft Certified: Azure Solutions Architect Expert, Google Professional Cloud Architect) or equivalent demonstrated experience.

Experience with FinOps practices and multi-cloud cost management tooling, including cost allocation and showback/chargeback models for shared research and administrative platforms.

Experience with enterprise private cloud and virtualization platforms (e.g., VMware vSphere) and integrating them with public cloud services in a hybrid mode.

Closing Date: 10/20/2026, 11:59PM ET Employee Group: USW Appointment Type : Budget - Continuing Schedule: Full-Time Pay Scale Group & Hiring Zone: USW Pay Band 19 -- $128,706 with an annual step progression to a maximum of $164,586.

Pay scale and job class assignment is subject to determination pursuant to the Job Evaluation/Pay Equity Maintenance Protocol.

Job Category: Information Technology (IT) Divisional HR Office Email: Lived Experience Statement Candidates who are members of Indigenous, Black, racialized and 2SLGBTQ+ communities, persons with disabilities, and other equity deserving groups are encouraged to apply, and their lived experience shall be taken into consideration as applicable to the posted position.

Diversity Statement The

University of Toronto embraces Diversity and is building aculture of belonging that increases our capacity to effectivelyaddress and serve the interests of our global community. Westrongly encourage applications from Indigenous Peoples,Black and racialized persons, women, persons withdisabilities, and people of diverse sexual and gender identities.We value applicants who have demonstrated a commitment toequity, diversity and inclusion and recognize that diverseperspectives, experiences, and expertise are essential tostrengthening our academic mission. As part of your application, you will be asked to complete a brief Diversity Survey.

This survey is voluntary.

Accessibility Statement The

University strives to be an equitable and inclusive community, and proactively seeks to increase diversity among its community members. Our values regarding equity and diversity are linked with our unwavering commitment to excellence in the pursuit of our academic mission. The University is committed to the principles of the Accessibility for Ontarians with Disabilities Act (AODA).

As such, we strive to make our recruitment, assessment and selection processes as accessible as possible and provide accommodations as required for applicants with disabilities.

📌 AI Infrastructure & Solutions Architect (Toronto)
🏢 University of Toronto
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai infrastructure & solutions architect (toronto) / toronto