The Lead Platform Reliability Engineer (PRE) ensures the stability, performance, and scalability of the shared platform that supports internal AI solution development. It combines software engineering, SRE practices, and operations to keep the platform reliable and developer-friendly.
Position Responsibilities
Reliability and performance: Define SLOs/SLIs, track operations budgets, reduce MTTR, capacity plan, and tune autoscaling.
Observability: Build and maintain logging, metrics, tracing, and alerting; instrument platform components; create runbooks and dashboards.
Incident response: On-call for platform incidents; triage, mitigate, root-cause, and drive postmortems and corrective actions.
Automation and tooling: Develop self-service capabilities, AIOps/MLOps/GitOps/CICD pipelines, and operational automations (provisioning, upgrades, backups).
Infrastructure as code: Manage clusters, networks, storage, and policies via Terraform/Ansible; prevent configuration drift.
Security and compliance: Enforce identity/RBAC, secrets management, supply chain security, and regulatory controls; collaborate with risk and audit.
Scalability and cost: Optimize resource usage, plan capacity, control spend (rightsizing, autoscaling, reservations/spot).
Change management: Safe rollouts, progressive delivery, and policy-as-code guardrails.
Platform productization: Treat the platform as a product, define operations SLAs in alignment to product roadmap, service catalog, and developer experience.
Collaborate with global engineering, security, and AI governance teams to ensure compliance with cross-geo regulations and Asia’s data residency requirements.
Operate scalable backend services supporting high-traffic agent interactions, retrieval operations, and real-time execution flows.
Maintain AI services runbooks, playbooks, and enablement for GOCC
Required Qualifications
Bachelor’s in Computer Science/Engineering or equivalent experience (not strictly required if skills demonstrated).
5-8 years experience in DevOps/Platform Engineering or Production Operations.
Proven track record operating large-scale distributed systems and running on-call.
Operational experience with cloud-native development: Azure, Kubernetes, containers, CI/CD, and observability stacks.
Knowledge with Python and/or Java/Scala/TypeScript for building backend services and automation.
Understanding of AI solution, LLM systems, retrieval architectures, embeddings, vector stores, prompt/tool orchestration, and agent workflow fundamentals.
Knowledge of API design, asynchronous workflows, concurrency, reliability engineering (SLOs, error budgets), and performance tuning.
Familiarity with security, governance, and compliance for AI/data systems (authN/authZ, data protection, audit logging, model governance).
Ability to collaborate across global teams and translate business requirements into platform capabilities and operational SLAs.
Preferred Qualifications
ITIL & ITSM certification
Azure Administrator/DevOps certificate (nice to have)
Kubernetes: CKA/CKS certificate (nice to have)
HashiCorp Terraform Associate certificate (nice to have)
When you join our team:
We’ll empower you to learn and grow the career you want.
We’ll recognize and support you in a flexible workplace where well-being and inclusion are more than just words.
As part of our global team, we’ll support you in shaping the future you want to see.
#LI-Hybrid
The role being advertised is an existing vacancy.
マニュライフとジョン・ハンコックについて
マニュライフ・ファイナンシャル・コーポレーションは、「あなたの未来に、わかりやすさを」を提供する、国際的な大手金融サービスプロバイダーです。当社について詳しくは、 https://www.manulife.co.jp/をご覧ください。
マニュライフは機会均等を是とする雇用主です
マニュライフ/ジョン・ハンコックでは、多様性を受け入れます。私たちは、サービス提供先であるお客さまと同様に、多様な人材を引きつけ、育成し、定着させ、文化や個人の力を受け入れる包括的な職場環境を促進するよう努めています。当社は公正な採用、定着、昇進、報酬に努めています。当社のすべての慣行およびプログラムは、人種、祖先、出身地、肌の色、民族的出身、市民権、宗教または宗教的信念、信条、性別(妊娠および妊娠関連の状態を含む)、性的指向、遺伝的特徴、退役軍人としての地位、性自認、性に関する表明、年齢、婚姻状況、家族状況、障害、または適用法で保護されるその他の要因に対する一切の差別を行うことなく管理されます。
雇用への平等なアクセスを提供するために、障壁を取り除くことが当社の優先事項です。人事担当者は、応募者が応募プロセス中に合理的配慮を要求する場合に協力します。配慮要求のプロセス中に共有されるすべての情報は、適用される法律およびマニュライフ/ジョン・ハンコックのポリシーに準拠した方法で保存および使用されます。申請プロセスにおいて合理的配慮を要求するには、
[email protected]までご連絡をお願いします。
Referenced Salary Location
Toronto, Ontario
Working Arrangement
ハイブリッド勤務
Salary range is expected to be between
$113,260.00 CAD - $210,340.00 CAD
Employees also have the opportunity to participate in incentive programs and earn incentive compensation tied to business and individual performance. The actual salary will vary depending on local market conditions, geography and relevant job-related factors such as knowledge, skills, qualifications, experience, and education/training. If you are applying for this role outside of the primary location, please contact
[email protected] for the salary range for your location.
Manulife offers eligible employees a wide array of customizable benefits, including health, dental, mental health, vision, short- and long-term disability, life and AD&D; insurance coverage, adoption/surrogacy and wellness benefits, and employee/family assistance plans. We also offer eligible employees various retirement savings plans (including pension and a global share ownership plan with employer matching contributions) and financial education and counseling resources. Our generous paid time off program in Canada includes holidays, vacation, personal, and sick days, and we offer the full range of statutory leaves of absence. If you are applying for this role in the U.S., please contact
[email protected] for more information about U.S.-specific paid time off provisions.
We use data and analytics technologies, such as artificial intelligence (AI), and automated processing tools, to analyze and process the information you provide to us or third parties in the application process. For more information, please refer to our personal information collection statement.
#J-18808-Ljbffr
📌 Lead Platform Reliability Engineer, Global AI Platform & Solutions (Ontario)
🏢 Manulife Financial
📍 Ontario