NYU Grossman trained a machine-learning screener on 14,555 past applications to flag interview-worthy candidates; a 2019 randomized trial of 3,700 applications found its recommendations statistically indistinguishable from faculty screeners overall, but it recommended economicall
14,555
Applications used to train the model (2013–2017)
6,000+ hours/year
Faculty screening hours previously required per year
2,910
Applications in retrospective validation
2,715
Applications in prospective validation (2018 cycle) (2018)
3,700
Applications in randomized trial (2019 cycle) (2019)
61/65 vs 70/71 (P=.14)
Underrepresented-in-medicine applicants recommended (algorithm vs faculty)
220/227 vs 224/229 (P=.55)
Female applicants recommended (algorithm vs faculty)
25/156 vs 44/178 (P=.048)
Economically disadvantaged applicants recommended (algorithm vs faculty)
Details
Maturity
Established
Promoter
NYU Grossman School of Medicine
Period
2013–2019
Keywords
admissions, higher education administration, machine learning screening, medical education
Context
NYU Grossman School of Medicine built a machine-learning tool to help screen the roughly 14,500 applications it receives each admissions cycle, a process that previously consumed more than 6,000 hours of faculty time a year. The model was trained on 14,555 applications from 2013–2017 and the human screening decisions faculty made on them.
Objectives
The tool aimed to reduce faculty screening workload while producing interview recommendations statistically comparable to human screeners, without replacing human judgment on final decisions.
Activities
The model went through staged validation: retrospective validation on 2,910 applications, prospective validation on 2,715 applications in the 2018 cycle, and finally a formal randomized trial in the 2019 cycle that split 3,700 applications between faculty-only screening and algorithm-assisted screening. It reads only structured data — not essays or letters — and its output remains a recommendation reviewed by humans, not a final decision.
Results
Published in the peer-reviewed journal Academic Medicine, the randomized trial found no statistically significant difference between the algorithm's and faculty's interview-recommendation rates overall, including for underrepresented-in-medicine applicants (61 of 65 vs 70 of 71 recommended, P=.14) and female applicants (220 of 227 vs 224 of 229, P=.55). The same trial found a statistically significant gap for economically disadvantaged applicants: the algorithm recommended 25 of 156 for interview versus 44 of 178 recommended by faculty (P=.048).
Conclusions
The authors are explicit that a model trained on past human decisions 'by definition' inherits the biases of the faculty who made them, that it only reads structured data, and that its output remains a recommendation reviewed by humans, not a final decision — the economically-disadvantaged-applicant gap is a genuine, disclosed equity caveat rather than a hidden flaw.
Implementation
Indicative cost
Medium (€50k–€500k)
Time to results
Medium (1–3 years)
Staffing & skills
NYU Grossman admissions faculty (screeners), Data science/machine-learning development team, Peer-reviewed publication team (Academic Medicine)
Conditions for success
A large historical labeled dataset of past applications and faculty decisions (14,555 records)
A staged validation process — retrospective, then prospective, then randomized trial — before adoption
Human review retained as the final decision-maker; the algorithm restricted to structured data only
Common failure modes
A model trained on past human decisions inherits historical faculty bias 'by definition' (the authors' own caveat)
Statistically significant under-recommendation of economically disadvantaged applicants (25/156 vs 44/178, P=.048)
Commonly funded by
Own resources / municipal budget
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
IDB/Elige Educar WhatsApp AI chatbot for career guidance to Chilean secondary students. Pre-registered RCT (43,000+ students, 2024–25): AI raised initiation …
Algeria's Ministry of Higher Education deployed a nationwide AI matching algorithm placing 340,901 baccalaureate graduates into university programmes in 2025 …