Staff Data Scientist, AI Evaluations Platform (Toronto)

Staff Data Scientist, AI Evaluations Platform (Toronto)

17 Aug
|
RBC
|
Toronto

17 Aug

RBC

Toronto

RBC’s AI Group is building trusted AI capabilities for the enterprise, and evaluation is one of the core controls that makes that possible.

As Staff Data

Scientist, AI Evaluations, you will be a senior technical leader in the data science function responsible for how RBC measures model and agent quality, safety, risk, and performance. You will design, build, and continuously improve evaluation datasets, sourcing methods, LLM judge design, deterministic scorers, human evaluation protocols, quality benchmarks, and measurement frameworks that help AI systems move from experimentation to production with evidence and control, while providing technical guidance and mentorship to other data scientists on the team. Design and drive advanced model and agent evaluation methodologies, including evaluation datasets, rubrics, LLM-as-judge methods, deterministic scorers, human evaluation, and measurement frameworks, providing technical guidance to junior team members.

Define evaluation science standards that translate model risk, responsible AI, product quality, safety, and business expectations into measurable criteria, repeatable methods, and clear evidence. Own the end-to-end lifecycle for evaluation datasets and scorecards, including sourcing, curation, validation, quality checks, versioning, lineage, reuse, and ongoing improvement. Design scalable evaluation approaches for generative AI and agentic systems, including task-level, workflow-level, trajectory-level, and runtime evaluation methods.

Partner with AI research, platform engineering, product, risk, governance, and business teams to embed evaluations into build, release, certification, monitoring, and recertification workflows. Establish human evaluation and review protocols that produce reliable labels, reviewer guidance, adjudication processes, quality controls, and audit-ready evidence.



Provide clear technical leadership, executive-ready communication, and mentorship to junior data scientists to help RBC scale trusted AI with speed, rigor, and control.

In this role, you will communicate and interact frequently with RBC partners and/or employees located across Canada and/or worldwide. 8+ years of experience in data science, applied machine learning, AI evaluation, ML quality, or a related technical field, including experience providing technical leadership and mentorship within a team. ~ Strong experience designing evaluation frameworks for ML, generative AI, or agentic AI systems, including metrics, datasets, benchmarks, rubrics, and quality measurement. ~ Practical experience with LLM evaluation methods such as LLM-as-judge, deterministic scoring, human evaluation, hallucination assessment, factuality assessment, safety evaluation, or model quality benchmarking. ~ Robust technical foundation in data science, statistics, machine learning, experimentation, data curation, Python, SQL, and modern AI/ML development practices. ~ Proven ability to translate governance, model risk, responsible AI, and business requirements into measurable controls, repeatable evaluation processes, and decision-ready evidence. ~ Strong communication and stakeholder management skills, with the ability to influence senior leaders across research, engineering, product, governance, risk, and business teams.

Experience evaluating agentic AI systems, tool-calling workflows,



multi-step reasoning, runtime traces, trajectory scoring, or workflow-level performance.

Experience in financial services, regulated AI, model risk management, responsible AI, enterprise governance, or audit-ready evidence processes. Familiarity with tools and platforms such as MLflow, Langfuse, LangSmith, OpenTelemetry, Grafana, CI/CD pipelines, or comparable evaluation and observability tooling. Publications, patents, open-source contributions, or industry work related to AI evaluation, ML quality, AI safety, applied research, or responsible AI.

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual. A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable A world-class training program in financial services Big Data Analytics, Critical Thinking, Decision Making, Industry Knowledge, Machine Learning (ML), Results-Oriented, Software Engineering, Software Product Design Employment Type: Full time Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities.

RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all. Join our Talent Community Find out how we use our passion and drive to enhance the well‑being of our clients and communities at jobs.rbc.com #

📌 Staff Data Scientist, AI Evaluations Platform (Toronto)
🏢 RBC
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: staff data scientist, ai evaluations platform (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: staff data scientist, ai evaluations platform (toronto) / toronto