evidoria

← Back to browse

Good practice Imported

Generative AI Without Guardrails Can Harm Learning — A Turkish High-School RCT of GPT Tutors

Türkiye · Ankara (exact school undisclosed for participant privacy) · See the Türkiye profile

Evidence: Randomised controlled trial Top 82% 33/100 · Ask Evidence Copilot about this practice

A randomised trial of nearly 1,000 Turkish high-schoolers found unrestricted GPT-4 ('GPT Base') lifted practice scores 48% but cut unassisted exam scores 17% versus controls; a pedagogically-guarded 'GPT Tutor' lifted practice scores 127% with no exam harm.

+48%
Grade improvement during assisted practice, GPT Base (Fall 2023-2024)
+127%
Grade improvement during assisted practice, GPT Tutor (Fall 2023-2024)
-17%
Unassisted exam score, GPT Base vs control (Fall 2023-2024)

Details

Promoter
Wharton School, University of Pennsylvania (Bastani et al. research team)
Period
Fall semester 2023–2024
Keywords
K-12 education, mathematics, generative AI, learning science

Context

In a large Turkish high school, close to 1,000 students in grades 9-11 were randomised into three arms across four 90-minute sessions during the Fall 2023-2024 semester: a no-AI control group, 'GPT Base' (unrestricted ChatGPT-style access), and 'GPT Tutor' (the same model with prompts designed to safeguard learning, such as withholding direct answers and prompting reasoning).

Results

During assisted practice, GPT Base improved grades by 48% and GPT Tutor by 127% relative to the control group. But on a later unassisted exam, GPT Base students scored 17% worse than controls, while GPT Tutor students, who had never been able to lean on direct answers, showed no significant harm.

Conclusions

Peer-reviewed in PNAS in 2025, the study is among the first large-scale randomised trials to show that generative AI can substitute for learning rather than support it when deployed without pedagogical guardrails — and that simple prompt-level safeguards can prevent that harm without sacrificing in-the-moment performance gains.

Implementation

Indicative cost
Low (< €50k) — Classroom-based research trial; no separate implementation budget is published.
Time to results
Short (< 1 year) — Four 90-minute sessions delivered during the Fall 2023-2024 semester.
Staffing & skills
Wharton School, University of Pennsylvania (Bastani et al. research team)

Conditions for success

  • Prompt-level safeguards (withholding direct answers, prompting reasoning) prevent the AI from substituting for learning

Common failure modes

  • Unrestricted chatbot access improved practice performance but harmed unassisted exam scores

Where it fits

Governance type
university research study
Scale
~1,000 students, one school
Income level
upper-middle income

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful