An exploratory RCT (N=165, 5 UK secondary schools) found human tutors supervising Google DeepMind's LearnLM boosted student learning gains to 10 percentage points versus 4.5 for human tutoring alone, with 76.4% of AI-drafted messages needing no more than minor edits.
10 pp
Learning gain vs standard hint, human-AI tutoring (Summer 2025)
4.5 pp
Learning gain vs standard hint, human tutoring alone (Summer 2025)
76.4 %
AI-drafted messages needing no more than minor edits
66.2 %
Novel-problem success rate, human-AI supported students
60.7 %
Novel-problem success rate, human-only tutored students
93.0 %
Immediate mistake-fix success rate, human-AI team
91.2 %
Immediate mistake-fix success rate, human tutor alone
95.4 %
Misconception-resolution success rate, human-AI team
94.9 %
Misconception-resolution success rate, human tutor alone
Details
Promoter
Eedi; Google DeepMind (LearnLM)
Period
Summer 2025 trial; published November 2025
Keywords
AI tutoring, mathematics education, human-AI collaboration, randomized controlled trial
Context
In summer 2025, Eedi and Google DeepMind ran an exploratory randomised controlled trial testing whether LearnLM, a generative AI model fine-tuned for pedagogy, could support human maths tutors without compromising instructional quality. The trial involved 165 students across five UK secondary schools on Eedi's chat-based maths tutoring platform, building on Eedi's earlier 2023-24 RCT of 2,901 students across 20 UK classrooms.
Objectives
The study set out to test whether AI-drafted tutoring messages, reviewed and editable by an expert human tutor before reaching the student, could match or improve on human-only tutoring, while keeping a human reviewer in control of every message sent.
Activities
In the human-AI condition, LearnLM drafted tutoring messages that an expert human tutor reviewed and could edit before they reached the student. Supervising tutors approved 76.4% of LearnLM's drafts with zero or only minor edits, and some AI-drafted Socratic questions reportedly taught tutors new pedagogical techniques.
Results
The human-AI team helped students fix an immediate mistake about as effectively as the human tutor alone (93.0% vs 91.2%) and resolve underlying misconceptions similarly well (95.4% vs 94.9%). The clearest advantage appeared on transfer: compared with a standard hint, human-only tutoring improved learning by 4.5 percentage points, while the human-AI team improved it by 10 percentage points, and students supported by LearnLM were more likely to solve novel problems on subsequent topics (66.2% vs 60.7%).
Conclusions
The authors describe the trial as exploratory and call for larger studies before broader conclusions are drawn: the sample of 165 students across 5 schools is small relative to the claims, and no equity-disaggregated results were reported.
Implementation
Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
Expert human maths tutors (Eedi), Eedi tutoring platform team, Google DeepMind LearnLM researchers
Conditions for success
Every AI-drafted message reviewed and editable by an expert human tutor before reaching the student
Use of an existing chat-based tutoring platform (Eedi) with established human tutors
Common failure modes
Small sample (165 students, 5 schools) limits generalisability
No equity-disaggregated results reported (by disadvantage, disability or other factors)
Commonly funded by
Own resources / municipal budget
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Replication kit
Reusable artefacts from this practice — as published by their sources.
A 13-week controlled study at four Hungarian universities found a ChatGPT-based adaptive tutor using Bayesian Knowledge Tracing nearly doubled programming …