Senior GenAI / LLM QA Engineer (Agentic AI Platforms) (Canada)

Senior GenAI / LLM QA Engineer (Agentic AI Platforms) (Canada)

09 Aug
|
Aqila Systems
|
Canada

09 Aug

Aqila Systems

Canada

We are seeking a highly experienced Senior GenAI / LLM QA Engineer to lead the quality assurance strategy and testing of a production-grade Agentic AI platform for the healthcare industry.

This is a senior, hands-on role requiring deep experience testing Generative AI applications, Large Language Model (LLM) integrations, AI agents, Retrieval-Augmented Generation (RAG), and multi-agent workflows operating in production environments.

The successful candidate will establish the AI testing framework, define evaluation methodologies, build automated testing capabilities, and ensure the platform consistently delivers reliable, accurate, secure, and safe AI-driven experiences suitable for healthcare organizations.

This role requires someone who has worked on production AI systems, not research prototypes, and understands the unique quality, reliability, and safety challenges associated with deploying LLM-powered applications at enterprise scale.

Key ResponsibilitiesAI Platform Quality Strategy

- Define and implement the overall QA strategy for an enterprise Agentic AI platform.
- Develop comprehensive testing methodologies for LLM-powered workflows, AI agents, orchestration engines, and RAG pipelines.
- Establish production-ready quality standards for AI-generated outputs, reasoning, workflow execution, and user interactions.

Functional & AI Testing

Design and execute test plans covering:

- Agentic AI workflows
- Multi-agent orchestration
- Prompt execution
- Tool calling and function execution
- RAG pipelines
- Knowledge retrieval accuracy
- AI decision-making
- Context management
- Memory handling
- Structured output validation
- Human-in-the-loop workflows

Validate:

- Accuracy
- Consistency
- Determinism where applicable
- Robustness
- Hallucination rates
- Response quality
- Grounding against source data
- Prompt adherence
- Safety guardrails
- Error recovery

Test Automation

Design and build automated testing frameworks for AI applications, including:

- Regression testing
- Prompt regression testing
- AI output validation
- API testing
- End-to-end workflow testing
- UI automation
- Performance testing
- Continuous AI evaluation
- CI/CD-integrated automated test suites

Develop reusable testing libraries and automation assets supporting continuous delivery.

Healthcare AI Validation

Ensure AI workflows meet healthcare-specific quality expectations including:

- Clinical workflow validation
- PHI handling and privacy controls
- Auditability
- Traceability
- Explainability
- Data integrity
- Safety testing
- Security validation





Collaborate with product, engineering, and clinical stakeholders to validate healthcare use cases and AI-assisted workflows.

AI Evaluation & Quality Metrics

Develop measurable evaluation frameworks for:

- Accuracy
- Precision
- Recall
- Groundedness
- Faithfulness
- Relevance
- Consistency
- Hallucination detection
- Response latency
- Workflow completion rates
- User acceptance metrics

Define production quality thresholds before release.

Collaboration

- Work closely with AI engineers, ML engineers, software developers, DevOps engineers, architects, and product teams.
- Participate in design reviews and AI architecture discussions.
- Identify production risks early and recommend mitigation strategies.
- Drive continuous quality improvements throughout the software lifecycle.

Required Qualifications

- 8+ years of experience in software quality assurance and automated testing.
- 4–5+ years of hands-on experience testing production Generative AI or LLM-powered applications.
- Demonstrated experience testing Agentic AI platforms deployed in production environments.
- Experience validating enterprise AI applications used by external customers.
- Solid understanding of Large Language Models and modern AI architectures.

Hands-on experience testing:

- Multi-agent systems
- AI orchestration frameworks
- RAG (Retrieval-Augmented Generation)
- Prompt engineering
- Function calling
- MCP (Model Context Protocol) integrations
- Vector databases and semantic search
- AI memory and context management

Experience building automated testing using tools such as:

- Playwright
- Cypress
- Selenium
- Postman
- Python
- JavaScript/TypeScript

Experience with API testing and automation.

Experience integrating automated testing into CI/CD pipelines.

Solid understanding of modern software development practices including Git and Agile methodologies.

Excellent analytical, debugging, and problem-solving skills.

Ability to work independently while collaborating effectively with cross-functional teams.

Preferred Qualifications

Experience working with one or more leading LLM platforms:

- OpenAI
- Anthropic Claude
- Google Gemini
- Azure OpenAI
- AWS Bedrock





Experience with AI orchestration frameworks such as:

- LangGraph
- LangChain
- CrewAI
- Semantic Kernel
- AutoGen

Experience evaluating AI systems using industry-standard frameworks and benchmarks.

Knowledge of prompt versioning, prompt optimization, and AI model evaluation methodologies.

Experience testing conversational AI, copilots, or enterprise AI assistants.

Healthcare industry experience, including familiarity with regulatory requirements such as HIPAA, PHIPA, or equivalent healthcare privacy frameworks.

Knowledge of AI safety, bias detection, adversarial testing, and responsible AI principles.

Experience working with Azure, AWS, or GCP AI platforms.

Engagement

The initial engagement will focus on establishing the AI quality assurance framework, automated testing strategy, evaluation methodology, and production validation processes for the healthcare Agentic AI platform.

The successful candidate will work closely with engineering and product teams throughout development and deployment, ensuring the platform meets enterprise standards for quality, reliability, security, and regulatory compliance. Ongoing involvement will include expanding automated test coverage, validating new AI capabilities, supporting production releases, and continuously improving AI quality as the platform evolves.

What We're Looking For

We are looking for a senior QA professional with proven experience testing production-grade Agentic AI and LLM-based applications. The ideal candidate has successfully built automated AI testing frameworks, understands the unique challenges of validating non-deterministic AI systems, and can ensure enterprise-grade quality for healthcare AI platforms.

The successful candidate will combine deep expertise in traditional software quality assurance with modern AI evaluation techniques to help deliver secure, reliable, and trustworthy AI solutions for healthcare organizations.

Pay: From $60.00 per hour

Application question(s):

- How many years of hands-on experience do you have testing production Generative AI or LLM-based applications?
- Have you tested production Agentic AI, AI Copilot, RAG, or multi-agent applications used by external customers?
- Do you have hands-on experience developing automated test frameworks for AI/LLM applications (e.g., prompt regression, API automation, UI automation, AI output validation)?

Experience:

- AI: 5 years (required)
- Healthcare: 5 years (preferred)

Work Location: Hybrid remote in Toronto, ON (Toronto District)

📌 Senior GenAI / LLM QA Engineer (Agentic AI Platforms) (Canada)
🏢 Aqila Systems
📍 Canada

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior genai / llm qa engineer (agentic ai platforms) (canada) / canada