A randomized controlled trial by Kyoto University added a GPT-4 chatbot co-facilitator to online discussions for 80 Afghan undergraduates (40 women, 40 men) barred from in-person study, raising posts per student from 5.11 to 8.73 (p<0.01) and self-reported confidence (p=0.037).
80 students (40 women, 40 men)
Participants
8.73 vs 5.11 posts/student
Posts per student (AI-facilitated vs control)
p=0.037
Self-reported post-session confidence
2.88 vs 0.78 likes/post
Peer 'likes' received (control vs treatment)
3,432 qualified applicants
Applicant pool
d>=0.65
Minimum detectable effect size
Details
Maturity
Pilot
Promoter
Kyoto University Graduate School of Informatics (Hyper-Democracy platform)
Period
February 2024 (published 2025–2026)
Keywords
AI in education, online learning, gender equity, conversational AI, conflict-affected education
Context
Afghanistan is the only country that systematically bans women and girls from secondary and higher education under Taliban rule since 2021; many Afghan women now rely on online learning provided by non-profits as their main route to education, but large online classes often struggle to sustain meaningful interaction.
Objectives
Researchers at Kyoto University's Graduate School of Informatics tested whether adding a GPT-4 chatbot co-facilitator to online post-lecture discussions could raise participation and confidence among Afghan students.
Activities
80 undergraduate computer science students (40 women, 40 men, aged 18-26) were recruited from a pool of 3,432 qualified applicants and randomly assigned to control or treatment groups. All students attended an identical 50-minute Zoom lecture, then joined 40-minute post-lecture discussions in gender-balanced groups of 10 on the open-source, Discourse-based 'Hyper-Democracy' platform. In the treatment condition, discussions were additionally co-facilitated by a GPT-4-based chatbot seeded with 17 human-authored facilitation prompts, alongside the same human instructor used in the control condition.
Results
Students in the AI-facilitated group posted significantly more messages (mean 8.73 vs 5.11 per student, p<0.01) and reported significantly higher post-session confidence in the material (p=0.037). Replies and word count were also higher but not statistically significant, and control-group posts received more peer 'likes' (2.88 vs 0.78, p<0.001) -- a pattern the authors attribute to a bandwagon effect concentrating likes on fewer control-group posts.
Conclusions
The trial was a single-day, single-course intervention, adequately powered only for large effects (Cohen's d >= 0.65), and the authors call for replication across longer courses and subjects before generalising. The work was funded by JSPS KAKENHI and JST CREST and published under a CC-BY-4.0 licence with a public data-access statement.
Implementation
Indicative cost
Low (< €50k) — Funded by JSPS KAKENHI and JST CREST (Japanese research grants); no cost figure disclosed.
Time to results
Short (< 1 year) — Single 50-minute lecture plus 40-minute discussion session, February 2024; published 2025-2026.
Staffing & skills
Kyoto University research team, human instructor (retained in both conditions), GPT-4 chatbot seeded with 17 human-authored facilitation prompts
Conditions for success
Gender-balanced small discussion groups (10 students, 5 women/5 men)