09 Aug
|
Aqila Systems
|
Canada
09 Aug
Aqila Systems
Canada
We are seeking a highly experienced Senior GenAI / LLM QA Engineer to lead the quality assurance strategy and testing of a production-grade Agentic AI platform for the healthcare industry.
This is a senior, hands-on role requiring deep experience testing Generative AI applications, Large Language Model (LLM) integrations, AI agents, Retrieval-Augmented Generation (RAG), and multi-agent workflows operating in production environments.
The successful candidate will establish the AI testing framework, define evaluation methodologies, build automated testing capabilities, and ensure the platform consistently delivers reliable, accurate, secure, and safe AI-driven experiences suitable for healthcare organizations.
This role requires someone who has worked on production AI systems, not research prototypes, and understands the unique quality, reliability, and safety challenges associated with deploying LLM-powered applications at enterprise scale.
Key ResponsibilitiesAI Platform Quality Strategy
- Define and implement the overall QA strategy for an enterprise Agentic AI platform.
- Develop comprehensive testing methodologies for LLM-powered workflows, AI agents, orchestration engines, and RAG pipelines.
- Establish production-ready quality standards for AI-generated outputs, reasoning, workflow execution, and user interactions.
Functional & AI Testing
Design and execute test plans covering:
- Agentic AI workflows
- Multi-agent orchestration
- Prompt execution
- Tool calling and function execution
- RAG pipelines
- Knowledge retrieval accuracy
- AI decision-making
- Context management
- Memory handling
- Structured output validation
- Human-in-the-loop workflows
Validate:
- Accuracy
- Consistency
- Determinism where applicable
- Robustness
- Hallucination rates
- Response quality
- Grounding against source data
- Prompt adherence
- Safety guardrails
- Error recovery
Test Automation
Design and build automated testing frameworks for AI applications, including:
- Regression testing
- Prompt regression testing
- AI output validation
- API testing
- End-to-end workflow testing
- UI automation
- Performance testing
- Continuous AI evaluation
- CI/CD-integrated automated test suites
Develop reusable testing libraries and automation assets supporting continuous delivery.
Healthcare AI Validation
Ensure AI workflows meet healthcare-specific quality expectations including:
- Clinical workflow validation
- PHI handling and privacy controls
- Auditability
- Traceability
- Explainability
- Data integrity
- Safety testing
- Security validation
Collaborate with product, engineering, and clinical stakeholders to validate healthcare use cases and AI-assisted workflows.
AI Evaluation & Quality Metrics
Develop measurable evaluation frameworks for:
- Accuracy
- Precision
- Recall
- Groundedness
- Faithfulness
- Relevance
- Consistency
- Hallucination detection
- Response latency
- Workflow completion rates
- User acceptance metrics
Define production quality thresholds before release.
Collaboration
- Work closely with AI engineers, ML engineers, software developers, DevOps engineers, architects, and product teams.
- Participate in design reviews and AI architecture discussions.
- Identify production risks early and recommend mitigation strategies.
- Drive continuous quality improvements throughout the software lifecycle.
Required Qualifications
- 8+ years of experience in software quality assurance and automated testing.
- 4–5+ years of hands-on experience testing production Generative AI or LLM-powered applications.
- Demonstrated experience testing Agentic AI platforms deployed in production environments.
- Experience validating enterprise AI applications used by external customers.
- Solid understanding of Large Language Models and modern AI architectures.
Hands-on experience testing:
- Multi-agent systems
- AI orchestration frameworks
- RAG (Retrieval-Augmented Generation)
- Prompt engineering
- Function calling
- MCP (Model Context Protocol) integrations
- Vector databases and semantic search
- AI memory and context management
Experience building automated testing using tools such as:
- Playwright
- Cypress
- Selenium
- Postman
- Python
- JavaScript/TypeScript
Experience with API testing and automation.
Experience integrating automated testing into CI/CD pipelines.
Solid understanding of modern software development practices including Git and Agile methodologies.
Excellent analytical, debugging, and problem-solving skills.
Ability to work independently while collaborating effectively with cross-functional teams.
Preferred Qualifications
Experience working with one or more leading LLM platforms:
- OpenAI
- Anthropic Claude
- Google Gemini
- Azure OpenAI
- AWS Bedrock
Experience with AI orchestration frameworks such as:
- LangGraph
- LangChain
- CrewAI
- Semantic Kernel
- AutoGen
Experience evaluating AI systems using industry-standard frameworks and benchmarks.
Knowledge of prompt versioning, prompt optimization, and AI model evaluation methodologies.
Experience testing conversational AI, copilots, or enterprise AI assistants.
Healthcare industry experience, including familiarity with regulatory requirements such as HIPAA, PHIPA, or equivalent healthcare privacy frameworks.
Knowledge of AI safety, bias detection, adversarial testing, and responsible AI principles.
Experience working with Azure, AWS, or GCP AI platforms.
Engagement
The initial engagement will focus on establishing the AI quality assurance framework, automated testing strategy, evaluation methodology, and production validation processes for the healthcare Agentic AI platform.
The successful candidate will work closely with engineering and product teams throughout development and deployment, ensuring the platform meets enterprise standards for quality, reliability, security, and regulatory compliance. Ongoing involvement will include expanding automated test coverage, validating new AI capabilities, supporting production releases, and continuously improving AI quality as the platform evolves.
What We're Looking For
We are looking for a senior QA professional with proven experience testing production-grade Agentic AI and LLM-based applications. The ideal candidate has successfully built automated AI testing frameworks, understands the unique challenges of validating non-deterministic AI systems, and can ensure enterprise-grade quality for healthcare AI platforms.
The successful candidate will combine deep expertise in traditional software quality assurance with modern AI evaluation techniques to help deliver secure, reliable, and trustworthy AI solutions for healthcare organizations.
Pay: From $60.00 per hour
Application question(s):
- How many years of hands-on experience do you have testing production Generative AI or LLM-based applications?
- Have you tested production Agentic AI, AI Copilot, RAG, or multi-agent applications used by external customers?
- Do you have hands-on experience developing automated test frameworks for AI/LLM applications (e.g., prompt regression, API automation, UI automation, AI output validation)?
Experience:
- AI: 5 years (required)
- Healthcare: 5 years (preferred)
Work Location: Hybrid remote in Toronto, ON (Toronto District)
📌 Senior GenAI / LLM QA Engineer (Agentic AI Platforms) (Canada)
🏢 Aqila Systems
📍 Canada