31 Aug
|
Appnovation
|
Toronto
31 Aug
Appnovation
Toronto
Who you are
- We are looking for people who can bring a strong, solution-focused mindset and contribute to quality standards, best practices and get things done
- Bachelor’s Degree in a technical field or equivalent experience
- 4+ years in QA / test engineering, with exposure to data/ML systems
- Strong Python and data-science techniques for measuring factual grounding and answer quality
- Experience with LLM evaluation frameworks and statistical analysis
- Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests
- Comfort working across multiple LLM providers’ outputs
- Test automation frameworks and scripting
- Detail-oriented, with strong analytical and communication skills
- You think about how to scale, automate and operate, not just how to build a solution to an immediate problem
- You understand lean thinking
- You set high standards for code quality, performance/scalability and security and seek continuous improvement
- You have solid analytical, problem solving and decision-making skills
- You have customer first mindset and a devotion to customer service
- You engage and build positive internal and external client relationships, while managing multiple initiatives, often with competing priorities
- You have strong self-initiative, passion, interpersonal, oral and written communication and collaboration skills with the ability to work, influence and make an impact in a cross-functional environment with all levels of the organization
- You are responsive and thrive in a fast-paced diverse high-performance setting with rapidly changing business needs
- You actively seek out things outside your comfort zone with the ability to rapidly learn and take advantage of new concepts, business models, and technologies
- You have prior experience in consulting
- Prior experience and connections in the Life Sciences industry is preferred
What the job involves
- As a QA / AI Evaluation Engineer, you will join a highly motivated and experienced team in a forward-leaning role that proves the platform actually improves answer quality. You will run evaluations at scale — from small human-UAT batches up to millions of automated evals — statistically measure factual grounding and accuracy lift, and build the metrics framework that shows how much better our answers get over time
- Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations
- Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing
- Build a metrics framework showing quality improvement (e.g., “answer is X% supported by source content / Y% better”)
- Design load and quality tests as the corpus scales
- Define and maintain test plans, test cases, and quality gates
- Automate regression and evaluation suites; integrate them into CI/CD
- Report quality metrics clearly to technical and non-technical stakeholders
- Collaborate with engineering to reproduce, triage, and verify fixes
- Continuously improve QA processes and coverage
📌 QA / AI Evaluation Engineer (Toronto)
🏢 Appnovation
📍 Toronto