10 Aug
|
Société Financière Manuvie
|
Toronto
10 Aug
Société Financière Manuvie
Toronto
The Lead Platform Reliability Engineer (PRE) ensures the stability, performance, and scalability of the shared platform that supports internal AI solution development. It combines software engineering, SRE practices, and operations to keep the platform reliable and developer-friendly.
Reliability and performance: Define SLOs/SLIs, track operations budgets, reduce MTTR, capacity plan, and tune autoscaling.
Automation and tooling: Develop self-service capabilities, AIOps/MLOps/GitOps/CICD pipelines, and operational automations (provisioning, upgrades, backups).
Infrastructure as code: Manage clusters, networks, storage, and policies via Terraform/Ansible; prevent configuration drift.
Security and compliance: Enforce identity/RBAC, secrets management, supply chain security, and regulatory controls; collaborate with risk and audit.
Scalability and cost: Optimize resource usage, plan capacity, control spend (rightsizing, autoscaling, reservations/spot).
Change management: Safe rollouts, progressive delivery, and policy-as-code guardrails.
Platform productization: Treat the platform as a product, define operations SLAs in alignment to product roadmap, service catalog, and developer experience. Collaborate with global engineering, security, and AI governance teams to ensure compliance with cross-geo regulations and Asia’s data residency requirements. Operate scalable backend services supporting high-traffic agent interactions, retrieval operations, and real-time execution flows.
Maintain AI services runbooks, playbooks, and enablement for GOCC Bachelor’s in Computer Science/Engineering or equivalent experience (not strictly required if skills demonstrated). ~5-8 years experience in DevOps/Platform Engineering or Production Operations. ~ Knowledge with Python and/or Java/Scala/TypeScript for building backend services and automation. ~ Understanding of AI solution, LLM systems,
retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflow fundamentals. ~ Knowledge of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning. ~ Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance). ~ Azure Administrator/DevOps certificate (nice to have) We’ll recognize and support you in a flexible workplace where well-being and inclusion are more than just words. As part of our global team, we’ll support you in shaping the future you want to see. #Le poste annoncé correspond à une vacance existante. 113,260.00 CAD - $210,340.00 CAD Les employés ont également la possibilité de participer à des programmes incitatifs et de recevoir une rémunération liée à la performance de l’entreprise et des individus. Le salaire réel variera selon les conditions du marché local, la région géographique et les facteurs propres au poste, tels que les connaissances, les compétences, les qualifications, l’expérience et la formation.
Manuvie offre aux employés admissibles une vaste gamme d’avantages sociaux personnalisables, notamment… une assurance soins médicaux soins dentaires assurance vie soins médicaux non urgents programmes d’aide aux employés et leur famille regimes d’épargne-retraite (y compris des régimes de rente et un programme international d’actionnariat assortie de cotisations patronales de contrepartie) ressources en matière d’éducation et de conseils financiers Nous utilisons des technologies de données et d’analytique, telles que l’intelligence artificielle (IA), ainsi que des outils de traitement automatisé pour analyser et traiter les renseignements que vous nous fournissez ou que des tiers nous transmettent dans le cadre du processus de demande.
📌 Lead platform reliability engineer, global ai platform & solutions (Toronto)
🏢 Société Financière Manuvie
📍 Toronto