evidoria

← Back to browse

Good practice Imported

AI-Assisted vs Human-Only Evidence Review — UK Government's Controlled Comparison Finds Speed Gains and Quality Trade-offs

United Kingdom · London · See the United Kingdom profile · See the London profile

Evidence: Quasi-experimental Top 24% 73/100 · Ask Evidence Copilot about this practice

DSIT and DCMS commissioned the Behavioural Insights Team to run matched AI-assisted and human-only rapid evidence reviews on the same brief. The AI-assisted review finished 23% faster but needed more fluency revisions, and the two shared only 4 of ~30 cited sources.

90.5 hours
Time to complete the AI-assisted review (2024)
117.75 hours
Time to complete the human-only review (2024)
23 %
Overall time saving with AI assistance (2024)
15 hours
Literature-analysis phase time, AI-assisted (2024)
34 hours
Literature-analysis phase time, human-only (2024)
56 %
Time reduction in the literature-analysis phase (2024)
4 of ~30 sources
Sources shared between the two reviews (2024)
AI-Assisted vs Human-Only Evidence Review — UK Government's Controlled Comparison Finds Speed Gains and Quality Trade-offs

Details

Maturity
Pilot
Promoter
Behavioural Insights Team (BIT), commissioned by the Department for Science, Innovation and Technology (DSIT) and the Department for Digital, Culture, Media and Sport (DCMS)
Period
2024
Keywords
evidence review, generative AI, policy analysis, government evaluation methods

Context

UK government departments increasingly consider using generative AI to speed up the rapid evidence reviews that inform policy advice, but had limited controlled evidence on how AI-assisted reviews compare with conventional human-only ones on the same brief.

Objectives

In 2024 the Department for Science, Innovation and Technology (DSIT) and the Department for Digital, Culture, Media and Sport (DCMS) commissioned the Behavioural Insights Team (BIT) to test whether generative AI can meaningfully speed up rapid evidence reviews without an unacceptable loss of quality.

Activities

Two separate researchers, given the same brief and inclusion criteria on how technology diffusion affects UK growth and productivity, independently produced a rapid evidence review, one working entirely by hand, the other using a mix of AI tools (including Elicit, Consensus, Claude 2 and ChatGPT-4) supplemented by manual checking and editing.

Results

The AI-assisted review was completed in 90.5 hours versus 117.75 hours for the human-only review, a 23% overall time saving, with the largest gap in the literature-analysis phase (15 hours versus 34 hours, a 56% reduction). However, the AI-assisted draft was judged less fluent and needed more revision, and the two reviews had surprisingly little overlap in the sources they cited, only 4 shared references out of roughly 30 unique sources found credible across both.

Conclusions

BIT and the commissioning departments concluded that generative AI can meaningfully speed up rapid evidence reviews but still produces errors requiring manual verification and can lead reviewers toward different bodies of evidence altogether; they recommended further work before wider adoption rather than treating the result as a mandate to automate evidence reviews.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year) — The comparison was conducted in 2024 as a one-off matched study, not an ongoing deployed practice.
Staffing & skills
Two independent researchers (one AI-assisted, one human-only), Behavioural Insights Team (BIT) as delivery and evaluation body, Department for Science, Innovation and Technology (DSIT) and Department for Digital, Culture, Media and Sport (DCMS) as commissioning departments

Conditions for success

  • Identical brief and inclusion criteria given to both reviewers to enable a fair comparison
  • Manual checking and editing of AI outputs rather than trusting them directly
  • Combining multiple AI tools (Elicit, Consensus, Claude 2, ChatGPT-4) rather than relying on a single tool

Common failure modes

  • The AI-assisted draft was judged less fluent and needed more revision than the human-only draft.
  • The two reviews shared only 4 of roughly 30 credible sources, showing AI assistance can steer reviewers toward a different evidence base.
  • The authors recommend further work before wider adoption rather than treating this as a mandate to automate evidence reviews.

Where it fits

Governance type
national government department, delivered via commissioned external evaluation
Scale
single matched study
Income level
high-income

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful