30 Jul
|
Deccan AI Experts
|
Canada
30 Jul
Deccan AI Experts
Canada
1.About the Project
Deccan AI is building a Deep Research Agent (DRA) Benchmark for the Legal domain as part of its GDPVal evaluation study. The benchmark evaluates how well frontier AI agents (OpenAI o3-deep-research, Google Gemini Deep Research, Claude) perform on realistic, high-difficulty qualified legal tasks-tasks where a standard frontier model fails but a genuine DRA can succeed.
The legal benchmark covers ten selected domains: Commercial Litigation, Corporate Law, Contract Review & Management, Regulatory Compliance, Intellectual Property, Consumer Protection, Employment & Labor Law, Environmental Law, Legal Research & Writing, and Policy & Legislative Analysis. Each sub-domain produces prompts graded by a two-layer scoring system: deterministic ground-truth verifiers and a five-criterion 0-3 SME rubric, composed into a Verifier-Rubric Score (VRS).
The Principal Domain Expert is the senior legal authority for this programme. This is not a task annotation role. The Principal is responsible for the intellectual integrity of the benchmark-from taxonomy design to SME screening to quality control of submitted prompt packages.
2.Role Summary
You will serve as the lead legal authority for the DRA Benchmark programme. You will own four inter-related workstreams: finalising the legal taxonomy and sub-domain framework; designing and validating the AI-driven SME interviewer used to screen and onboard task-writing legal experts; calibrating the task complexity and scoring framework (Numerical Rigor Levels L0-L5, prompt types CRP/RCP/SCP/LDP/FSP, decision archetypes); and conducting QC reviews of submitted prompt packages to ensure they meet the benchmark's analytical rigour, legal reasoning, and verifiability standards.
This role requires someone who can simultaneously think like a senior legal practitioner and an evaluation scientist. You must be able to identify what makes a legal task genuinely difficult for an AI agent, design structures that make failure measurable, and ensure that legal SMEs apply those standards consistently.
3.Key Responsibilities
3.1 Taxonomy Finalisation
- Review and validate the selected legal domains and formally document the rationale for inclusion and exclusion decisions.
- Define precise boundaries between overlapping legal domains, including:
- Commercial Litigation (e.g., Litigation Risk Assessment and Strategy ) vs. Corporate Law (e.g., Board Advisory and Compliance )
- Regulatory Compliance (e.g., Industry-Specific Regulatory Frameworks (sectoral licensing, regulated industries, permitting regimes, etc.) ) vs. Consumer Protection (e.g.,
Unfair and Deceptive Trade Practices )
- Employment & Labor Law (e.g., Workplace Safety and Occupational Safety Regulations ) vs. Regulatory Compliance (e.g., Administrative Hearings and Agency Proceedings )
- Legal Research & Writing (e.g., Statutory and Regulatory Research ) vs. Policy & Legislative Analysis (e.g., Regulatory Impact Analysis )
- Contract Review & Management (e.g., IP Licensing Agreements overlap) vs. Intellectual Property (e.g., IP Licensing Agreements, IP Litigation and Settlement )
- Build and maintain a cross-domain persona capability matrix indicating which legal personas can author or review tasks across domains. Example personas:
- Commercial Litigation Counsel
- Corporate Counsel
- Contracts Manager
- Compliance Officer
- IP Attorney
- Employment Law Specialist
- Environmental Counsel
- Policy Analyst
- Validate and correct domain tags on all reference tasks and ensure consistent labeling across benchmark documentation.
3.2 AI Interviewer Design for SME Onboarding
- Design the AI-driven screening interview used to evaluate incoming legal SMEs before onboarding.
- Define competency dimensions the interview must assess:
- Legal technical depth
- Statutory and regulatory interpretation
- Case-law, precedent, and legal authority analysis
- Ability to identify likely AI reasoning failures
- Ability to write objective, binary-checkable verifiers
- Legal drafting and analytical reasoning skills
- Author or review the screening question bank.
- Define pass/fail thresholds for each legal sub-domain and persona level.
- Determine when borderline candidates require secondary human review.
- Test and calibrate the AI interviewer against expert legal judgment before deployment.
3.3 Task Complexity Framework Calibration
- Validate the Numerical Rigor Level (L0-L5) framework and identify misclassifications.
- Finalize legal prompt-type definitions and cognitive trap frameworks.
- Define complexity indicators such as:
- Multi-jurisdictional analysis
- Conflicting precedents
- Regulatory ambiguity
- Statutory interpretation challenges
- Contractual conflict resolution
- Litigation strategy trade-offs
- Policy impact analysis
- Ensure complexity levels accurately reflect professional legal practice rather than simple legal knowledge retrieval.
3.4 Quality Control (QC) of Submitted Prompt Packages
- Review submitted prompt packages against the four-component standard:
- Prompt
- Data Files
- Solution Logic
- Sanity Check
- Verify solution logic step-by-step and independently validate legal conclusions.
- Confirm legal authorities, statutes, regulations, precedents, and contractual provisions cited in solutions.
- Validate prompt-specific verifiers to ensure they are:
- Binary-checkable
- Legally grounded
- Derived from the expected solution
- Non-redundant with rubric criteria
- Issue ACCEPT / REVISE / REJECT verdicts with written justification.
- Maintain a QC log of recurring legal reasoning and task-design issues.
- Escalate benchmark calibration issues where tasks are systematically too easy or too difficult for evaluated models.
Required Qualifications
- Professional legal qualification required (e.g., J.D., LL.B., B.L., or equivalent law degree recognized in the candidate's jurisdiction).
- Minimum 10 years of post-qualification legal experience in law firms, corporate legal departments, regulatory bodies, consulting firms, policy organizations, or litigation practice.
- Experience working with, supervising, reviewing, or evaluating legal work products across multiple jurisdictions is strongly preferred. Candidates should be comfortable assessing legal reasoning, legal research, and legal work products originating from both United States and Indian legal contexts.
- Demonstrable expertise across at least three of the benchmark domains:
- Commercial Litigation
- Corporate Law
- Contract Review & Management
- Regulatory Compliance
- Intellectual Property
- Consumer Protection
- Employment & Labor Law
- Environmental Law
- Legal Research & Writing
- Policy & Legislative Analysis
- Strong familiarity with:
- Statutory interpretation
- Regulatory frameworks
- Contract drafting and review
- Legal research methodologies
- Judicial precedent analysis
- Legal writing standards
Preferred Qualifications
- Professional governance, corporate secretarial, compliance, or corporate governance qualifications/experience preferred.
- Prior experience in AI evaluation, benchmark design, RLHF annotation, legal-tech products, or legal knowledge systems.
- Experience designing interviews, assessments, or competency frameworks for legal professionals.
- Cross-sector experience in litigation, corporate advisory, compliance, technology law, public policy, or regulatory affairs.
📌 Legal Domain Expert (Remote | Freelance) (Canada)
🏢 Deccan AI Experts
📍 Canada