evidoria

← Back to browse

Good practice Imported

AHEAD Trial — AI-Generated Medical Exam Questions Match Human-Written Ones on Most Measures, With a Flagged Equity Gap

Canada · Vancouver · See the Canada profile · See the Vancouver profile

Top 59% 47/100 · Ask Evidence Copilot about this practice

A blinded RCT of 258 first-year UBC medical students found AI-generated exam questions were made 5.6x faster than student-written ones, with similar scores and acceptability, though human items discriminated slightly better and a gender gap emerged only on the AI exam.

AHEAD Trial — AI-Generated Medical Exam Questions Match Human-Written Ones on Most Measures, With a Flagged Equity Gap

Details

Promoter
University of British Columbia Faculty of Medicine
Period
December 2024 (published June 2026)
Keywords
higher education, medical education, assessment design, academic integrity

Description

Writing a single high-quality multiple-choice exam question for medical education can take instructors or item-writers well over an hour and cost an estimated US$200 to over US$2,000 per usable item, making large question banks expensive to build and maintain.
The AHEAD Trial, run by the University of British Columbia Faculty of Medicine on 8-9 December 2024, randomized 258 first-year MD students 1:1 to a blinded 112-item mock exam built either by senior medical students (human-generated) or by a ChatGPT-4-plus-Google-Gemini validation pipeline (AI-generated), with all authorship indicators stripped before delivery.
AI-generated items took 4.2 minutes each to produce versus 19.6 minutes for human-written items — a 5.6-fold efficiency gain (p<0.0001). Overall exam scores were statistically similar (73.4% human-generated vs 71.0% AI-generated, p=0.083), and student-rated acceptability differed only slightly, favoring the human exam on four of nine measures with small effect sizes (Cohen's d ≤0.32). Human-written items had modestly but significantly higher item-discrimination scores than AI-generated ones (small effect), and neither exam meaningfully changed students' self-rated exam readiness.
An exploratory, likely underpowered subgroup analysis found male students outperforming female students on the AI-generated exam (rrb=0.26) but not on the human-generated exam (rrb=0.07) — a signal the authors flag as requiring demographic monitoring rather than a confirmed effect. The trial was single-center, involved only first-year students, question-writers were senior students rather than faculty, and it was registered retrospectively, all of which the authors list as limitations.

Read the full analysis: https://zenodo.org/records/18284890

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful