evidoria

← Back to browse

Good practice

CDA-AMC's Evaluation of AI Search Tools for Health-Technology-Assessment Evidence Synthesis

Canada · Ottawa · See the Canada profile

Canada's Drug Agency ran a peer-reviewed, 7-project comparison of Lens.org, SpiderCite and Microsoft Copilot against manual systematic-review search for drug-reimbursement evidence, finding all three AI tools recalled far fewer relevant studies than trained searchers.

0.986
Recall/sensitivity, manual search (7 HTA projects, Aug-Nov 2024)
0.676
Recall/sensitivity, Lens.org (7 HTA projects, Aug-Nov 2024)
0.23-0.26
Recall/sensitivity, SpiderCite (7 HTA projects, Aug-Nov 2024)
0.244
Recall/sensitivity, Microsoft Copilot (GPT-4) (7 HTA projects, Aug-Nov 2024)
2.88 hours
Average search time, manual
0.96 hours
Minimum average search time, Copilot

Details

Promoter
Canada's Drug Agency (CDA-AMC, formerly CADTH)
Period
2024–2025
Keywords
evidence synthesis, health technology assessment, AI tool evaluation, regulatory policy analysis, systematic review

Context

Canada's Drug Agency (CDA-AMC, formerly CADTH) tasked its Research Information Services team with evaluating three AI-assisted literature-search and evidence-synthesis tools — Lens.org, SpiderCite and Microsoft Copilot (GPT-4) — against standard systematic-review search methods, for the Health Technology Assessment (HTA) work underpinning Canada's drug and technology reimbursement recommendations.

Objectives

To assess whether AI-assisted search tools can match manual systematic-review search performance for HTA evidence synthesis.

Activities

A retrospective comparison across 7 completed agency HTA projects (August-November 2024) benchmarked each AI tool's search recall and search time against manual reference search. Findings were published as a peer-reviewed comparative study and incorporated into an official CDA-AMC position statement on AI use in HTA.

Results

Manual reference search achieved a recall/sensitivity of 0.986, versus 0.676 for Lens.org, roughly 0.23-0.26 for SpiderCite, and 0.244 for Copilot. Search time fell from an average of 2.88 hours for manual search to as little as 0.96 hours for Copilot.

Conclusions

The agency's conclusion is cautionary rather than promotional: each tool was judged 'fit for purpose' only in narrow supporting roles — for example, Copilot for search-strategy brainstorming — and explicitly not as a replacement for standard systematic search on complex HTA topics.

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful