Ai Engineer Sr (Winnipeg)

Ai Engineer Sr (Winnipeg)

09 Oct
|
Dayforce
|
Winnipeg

09 Oct

Dayforce

Winnipeg

About the opportunityWe are looking for an AI Engineer - Agentic Systems Evaluation to help define how we measure, test, and improve the quality of enterprise AI systems.
As AI evolves from conversational assistants and Retrieval-Augmented Generation (RAG) into tool using agents, multi-agent systems, and autonomous business workflows, evaluating only the final response is no longer enough.
An agent may reach the right answer while choosing the wrong tool, taking unnecessary steps, retrieving incorrect context, failing to escape to a human or violating business rules.
This role will build the evaluation frameworks needed to understand not only whether an AI system succeeded, but how it succeeded, how reliably it can repeat that outcome, and whether the architecture is appropriate for the problem.
What you’ll get to doBuild Agentic Evaluation Frameworks
Design evaluation methodologies covering the complete AI execution lifecycle:
Intent Planning Retrieval Tool Use Reasoning Action Business Outcome
Evaluate systems across dimensions including:
task and business outcome accuracyplanning and decision qualityretrieval quality and groundednesstool selection and executionagent routing and delegationhuman escalation decisionsreliability and failure recoverysafety and policy compliancelatency, token consumption, and cost
Evaluate Agentic Design Patterns
Design experiments and benchmarks that help engineering teams determine which architecture works best for a given problem.
Evaluate patterns such as:
Single-agent vs. multi-agent systemsSupervisor/router architecturesPlanner–executor patternsSequential and parallel workflowsTool-using agentsHuman-in-the-loop workflowsLong-running agents with state and memory
Measure whether additional agent complexity actually improves task success, reliability, and business outcomes enough to justify increased latency, cost, and operational complexity.
Advance RAG & Knowledge Evaluation




Build rigorous evaluation approaches for enterprise retrieval and knowledge systems, including:
retrieval precision and relevancegroundedness and citation accuracysource authority and freshnesschunking, metadata, and indexing strategiessemantic vs. hybrid search and rerankingpermission-aware retrieval
Evaluate how retrieval decisions ultimately impact downstream agent performance, rather than treating RAG evaluation as an isolated problem.
Build Automated Evaluation & Regression Testing
Develop scalable evaluation infrastructure including:
golden and synthetic datasetsscenario and adversarial test suitesdeterministic gradersLLM-as-a-Judge evaluationhuman evaluation workflowstrace-based evaluationautomated regression testing
Integrate evaluations into AI development and release pipelines so changes to models, prompts, retrieval, tools, or agent architectures can be measured before reaching production.
Build Agent Trace & Failure Analysis
Analyze complete agent execution traces including planning, retrieved context, tool calls, handoffs, retries, exceptions, latency, and cost.
Develop failure taxonomies that distinguish between:
Model | Retrieval | Planning | Tool | Routing | Memory | Integration | Policy | Orchestration failures
Turn production failures and user feedback into measurable regression tests and engineering improvements.
Define Production AI Quality
Establish measurable quality standards and release criteria for AI systems.
Metrics may include Task Success Rate, First-Pass Success Rate, Tool Selection Accuracy, Agent Routing Accuracy, Plan Execution Fidelity, Failure Recovery Rate, Human Escalation Accuracy,



Business Outcome Accuracy, Cost / Latency per Successful Task
Help teams answer a fundamental question:
Is this AI system reliable, secure, efficient, and valuable enough to operate in production?
Skills and experience we valueStrong experience with:
Python and software engineeringLLM application developmentAI/ML evaluation and experimentationautomated testing and data analysisRAG, embeddings, hybrid search, and rerankingtool/function calling and structured outputsagent orchestration and multi-agent workflowsAI observability and tracingAPIs and enterprise integrations
Experience with platforms or frameworks such as Agents SDK, LangGraph/LangChain, Microsoft AI Foundry, Amazon Bedrock, or similar agent platforms is valuable.
You should be comfortable working with evaluation techniques such as offline/online evals, deterministic graders, model-based graders, human evaluation, synthetic datasets, adversarial testing.
What would make you stand outYou don't stop when an agent successfully completes a task. You ask:
Did it choose the right approach?Did it use the right tools and information?Can it succeed consistently?Can it recover when something fails?Did adding more agents improve the outcome?Could a simpler architecture achieve the same result?What did the successful outcome cost?
You turn those questions into measurable experiments that help engineering teams build better AI systems.
Why This Role MattersEnterprise AI is moving from systems that answer questions to systems that make decisions and perform work.
That changes how quality must be measured.
The next generation of AI evaluation must measure:
What the system understood what it retrieved what it decided what it did and whether the business outcome was correct.
This role will help establish the engineering discipline required to make those systems measurable, reliable, and production-ready.
#J-18808-Ljbffr

📌 Ai Engineer Sr (Winnipeg)
🏢 Dayforce
📍 Winnipeg

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai engineer sr (winnipeg) / winnipeg

Subscribe to this job alert:

Get the latest job offers by email for: ai engineer sr (winnipeg) / winnipeg