23 Aug
|
United States Digital Space
|
Kitchener
23 Aug
United States Digital Space
Kitchener
the company is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.
Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, the company was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.
At the company, AI isn’t just a feature; We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.
And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.
We look for people who are intensely curious and hold themselves to a high bar. As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for the company’s Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.
This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.
You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.
You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams.
You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis.
Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
Experience designing structured test strategies across manual and automated workflows.
Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
Robust analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.
📌 AI Evaluation Engineer (Kitchener)
🏢 United States Digital Space
📍 Kitchener