A blinded RCT of 258 first-year UBC medical students found AI-generated exam questions were made 5.6x faster than student-written ones, with similar scores and acceptability, though human items discriminated slightly better and a gender gap emerged only on the AI exam.
Details
Promoter
University of British Columbia Faculty of Medicine
Period
December 2024 (published June 2026)
Keywords
higher education, medical education, assessment design, academic integrity
Description
Writing a single high-quality multiple-choice exam question for medical education can take instructors or item-writers well over an hour and cost an estimated US$200 to over US$2,000 per usable item, making large question banks expensive to build and maintain. The AHEAD Trial, run by the University of British Columbia Faculty of Medicine on 8-9 December 2024, randomized 258 first-year MD students 1:1 to a blinded 112-item mock exam built either by senior medical students (human-generated) or by a ChatGPT-4-plus-Google-Gemini validation pipeline (AI-generated), with all authorship indicators stripped before delivery. AI-generated items took 4.2 minutes each to produce versus 19.6 minutes for human-written items — a 5.6-fold efficiency gain (p<0.0001). Overall exam scores were statistically similar (73.4% human-generated vs 71.0% AI-generated, p=0.083), and student-rated acceptability differed only slightly, favoring the human exam on four of nine measures with small effect sizes (Cohen's d ≤0.32). Human-written items had modestly but significantly higher item-discrimination scores than AI-generated ones (small effect), and neither exam meaningfully changed students' self-rated exam readiness. An exploratory, likely underpowered subgroup analysis found male students outperforming female students on the AI-generated exam (rrb=0.26) but not on the human-generated exam (rrb=0.07) — a signal the authors flag as requiring demographic monitoring rather than a confirmed effect. The trial was single-center, involved only first-year students, question-writers were senior students rather than faculty, and it was registered retrospectively, all of which the authors list as limitations.
Read the full analysis: https://zenodo.org/records/18284890
Implementation
Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.