evidoria

← Back to browse

Good practice Imported

Eedi × Google DeepMind LearnLM — Human-in-the-Loop AI Tutoring RCT in UK Secondary Schools

United Kingdom · London · See the United Kingdom profile · See the London profile

Evidence: Randomised controlled trial Top 28% 60/100 · Ask Evidence Copilot about this practice

An exploratory RCT (N=165, 5 UK secondary schools) found human tutors supervising Google DeepMind's LearnLM boosted student learning gains to 10 percentage points versus 4.5 for human tutoring alone, with 76.4% of AI-drafted messages needing no more than minor edits.

10 pp
Learning gain vs standard hint, human-AI tutoring (Summer 2025)
4.5 pp
Learning gain vs standard hint, human tutoring alone (Summer 2025)
76.4 %
AI-drafted messages needing no more than minor edits
66.2 %
Novel-problem success rate, human-AI supported students
60.7 %
Novel-problem success rate, human-only tutored students
93.0 %
Immediate mistake-fix success rate, human-AI team
91.2 %
Immediate mistake-fix success rate, human tutor alone
95.4 %
Misconception-resolution success rate, human-AI team
94.9 %
Misconception-resolution success rate, human tutor alone
Eedi × Google DeepMind LearnLM — Human-in-the-Loop AI Tutoring RCT in UK Secondary Schools

Details

Promoter
Eedi; Google DeepMind (LearnLM)
Period
Summer 2025 trial; published November 2025
Keywords
AI tutoring, mathematics education, human-AI collaboration, randomized controlled trial

Context

In summer 2025, Eedi and Google DeepMind ran an exploratory randomised controlled trial testing whether LearnLM, a generative AI model fine-tuned for pedagogy, could support human maths tutors without compromising instructional quality. The trial involved 165 students across five UK secondary schools on Eedi's chat-based maths tutoring platform, building on Eedi's earlier 2023-24 RCT of 2,901 students across 20 UK classrooms.

Objectives

The study set out to test whether AI-drafted tutoring messages, reviewed and editable by an expert human tutor before reaching the student, could match or improve on human-only tutoring, while keeping a human reviewer in control of every message sent.

Activities

In the human-AI condition, LearnLM drafted tutoring messages that an expert human tutor reviewed and could edit before they reached the student. Supervising tutors approved 76.4% of LearnLM's drafts with zero or only minor edits, and some AI-drafted Socratic questions reportedly taught tutors new pedagogical techniques.

Results

The human-AI team helped students fix an immediate mistake about as effectively as the human tutor alone (93.0% vs 91.2%) and resolve underlying misconceptions similarly well (95.4% vs 94.9%). The clearest advantage appeared on transfer: compared with a standard hint, human-only tutoring improved learning by 4.5 percentage points, while the human-AI team improved it by 10 percentage points, and students supported by LearnLM were more likely to solve novel problems on subsequent topics (66.2% vs 60.7%).

Conclusions

The authors describe the trial as exploratory and call for larger studies before broader conclusions are drawn: the sample of 165 students across 5 schools is small relative to the claims, and no equity-disaggregated results were reported.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
Expert human maths tutors (Eedi), Eedi tutoring platform team, Google DeepMind LearnLM researchers

Conditions for success

  • Every AI-drafted message reviewed and editable by an expert human tutor before reaching the student
  • Use of an existing chat-based tutoring platform (Eedi) with established human tutors

Common failure modes

  • Small sample (165 students, 5 schools) limits generalisability
  • No equity-disaggregated results reported (by disadvantage, disability or other factors)

Commonly funded by

Own resources / municipal budget

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful