29 Aug
|
Equifax
|
Toronto
At Equifax, we are moving past passive AI chat interfaces to build the future of autonomous workflows. We are creating intelligent, self‑correcting multi‑agent systems that can navigate complex software environments, utilize external tools, and solve open-ended business problems with minimal human intervention. To ensure these systems are secure, reliable, and enterprise‑grade, we are seeking an analytical Agentic AI Evaluation & Tuning Engineer.
In this role, you will be the guardian of our production AI reliability. You will bridge the gap between raw Large Language Model (LLM) capabilities and flawless autonomous execution. Unlike traditional software testers or prompt engineers, you will focus on the behavior, decision‑making logic, tool‑use efficiency, and long‑term stability of multi‑agent architectures.
Build the "Golden Set": Curate, maintain, and augment high‑quality reference datasets (Golden Sets) of documents, user queries, and expected agent trajectories to serve as the ultimate source of truth for testing.
Automate Eval Cycles: Design and implement automated, continuous evaluation pipelines to measure agent accuracy, latency, token spend, and fallback reliability before code hits production. ReAct, Reflection loops) to pinpoint exactly where an agent deviates from its intended logic path. Refine system prompts, context windows, and few‑shot examples to optimize how agents execute complex, multi‑step workflows.
Fine‑tune how agents interact with external APIs, databases, and UiPath RPA workflows—minimizing execution errors, redundant calls, and token overhead.
Partner closely with AI Solution Leads and AI Agent Developers to feed evaluation insights back into the development lifecycle, helping them build robust, reusable, and self‑correcting agent components. Implement robust guardrail frameworks to ensure agents maintain reliable, fact‑based autonomous decision‑making post‑deployment in production.
Optimize
Domain‑Specific Knowledge Bases and Retrieval‑Augmented Generation (RAG) pipelines to ensure agents pull from accurate data rather than assumptions.
Experience: 3+ years of professional experience in software quality engineering, test automation, or data/ML engineering, with a dedicated focus on LLM testing, prompt tuning, or orchestration patterns over the last 1–2 years.
Agentic & LLM Frameworks: Proven hands‑on experience working with LLM orchestration frameworks (e.g., Advanced Debugging & Automation: Strong background in writing automated test scripts (Python‑heavy) and using tracing/observability concepts to debug cascading errors in asynchronous, non‑deterministic systems.
Experience with AI evaluation and observability platforms Live production experience testing Agentic workflows and GenAI solutions Familiarity with Google Cloud AI suite (Vertex & Gemini Enterprise Agent Platform) and UiPath ecosystem (Maestro).
Experience utilizing LLMs to securely generate high‑quality synthetic data for edge‑case testing. Proficiency in Python or TypeScript, with a deep understanding of asynchronous programming, API design, and microservices architecture.
📌 Agentic AI Optimization Developer (Toronto)
🏢 Equifax
📍 Toronto