evidoria

← Back to browse

Good practice Imported

NYU Grossman School of Medicine's Machine-Learning Screener for Medical-School Applications

United States of America · New York · See the United States of America profile · See the New York profile

Evidence: Randomised controlled trial Top 77% 40/100 · Ask Evidence Copilot about this practice

NYU Grossman trained a machine-learning screener on 14,555 past applications to flag interview-worthy candidates; a 2019 randomized trial of 3,700 applications found its recommendations statistically indistinguishable from faculty screeners overall, but it recommended economicall

14,555
Applications used to train the model (2013–2017)
6,000+ hours/year
Faculty screening hours previously required per year
2,910
Applications in retrospective validation
2,715
Applications in prospective validation (2018 cycle) (2018)
3,700
Applications in randomized trial (2019 cycle) (2019)
61/65 vs 70/71 (P=.14)
Underrepresented-in-medicine applicants recommended (algorithm vs faculty)
220/227 vs 224/229 (P=.55)
Female applicants recommended (algorithm vs faculty)
25/156 vs 44/178 (P=.048)
Economically disadvantaged applicants recommended (algorithm vs faculty)

Details

Maturity
Established
Promoter
NYU Grossman School of Medicine
Period
2013–2019
Keywords
admissions, higher education administration, machine learning screening, medical education

Context

NYU Grossman School of Medicine built a machine-learning tool to help screen the roughly 14,500 applications it receives each admissions cycle, a process that previously consumed more than 6,000 hours of faculty time a year. The model was trained on 14,555 applications from 2013–2017 and the human screening decisions faculty made on them.

Objectives

The tool aimed to reduce faculty screening workload while producing interview recommendations statistically comparable to human screeners, without replacing human judgment on final decisions.

Activities

The model went through staged validation: retrospective validation on 2,910 applications, prospective validation on 2,715 applications in the 2018 cycle, and finally a formal randomized trial in the 2019 cycle that split 3,700 applications between faculty-only screening and algorithm-assisted screening. It reads only structured data — not essays or letters — and its output remains a recommendation reviewed by humans, not a final decision.

Results

Published in the peer-reviewed journal Academic Medicine, the randomized trial found no statistically significant difference between the algorithm's and faculty's interview-recommendation rates overall, including for underrepresented-in-medicine applicants (61 of 65 vs 70 of 71 recommended, P=.14) and female applicants (220 of 227 vs 224 of 229, P=.55). The same trial found a statistically significant gap for economically disadvantaged applicants: the algorithm recommended 25 of 156 for interview versus 44 of 178 recommended by faculty (P=.048).

Conclusions

The authors are explicit that a model trained on past human decisions 'by definition' inherits the biases of the faculty who made them, that it only reads structured data, and that its output remains a recommendation reviewed by humans, not a final decision — the economically-disadvantaged-applicant gap is a genuine, disclosed equity caveat rather than a hidden flaw.

Implementation

Indicative cost
Medium (€50k–€500k)
Time to results
Medium (1–3 years)
Staffing & skills
NYU Grossman admissions faculty (screeners), Data science/machine-learning development team, Peer-reviewed publication team (Academic Medicine)

Conditions for success

  • A large historical labeled dataset of past applications and faculty decisions (14,555 records)
  • A staged validation process — retrospective, then prospective, then randomized trial — before adoption
  • Human review retained as the final decision-maker; the algorithm restricted to structured data only

Common failure modes

  • A model trained on past human decisions inherits historical faculty bias 'by definition' (the authors' own caveat)
  • Statistically significant under-recommendation of economically disadvantaged applicants (25/156 vs 44/178, P=.048)

Commonly funded by

Own resources / municipal budget

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful