evidoria

← Back to browse

Good practice Imported

AI-Assisted Oral Reading Screening — Flagging Reading Difficulty in Greek Primary Pupils

Greece · Volos · See the Greece profile

Evidence: Observational / pre–post Top 59% 47/100 · Ask Evidence Copilot about this practice

University of Thessaly researchers built a deep-learning tool that flags reading difficulty from 7-second recordings of Greek pupils reading aloud, reaching 84% accuracy vs expert judgement — a low-cost screening aid the authors say still needs classroom-noise validation.

77 pupils
Validation sample size (study sample, grades 3-6)
84 %
Overall accuracy vs expert assessment (study)
0.85 score
Balanced accuracy (study)
0.89 score
Macro F1 score (study)
0.74 correlation coefficient
Correlation with expert judgment (study)

Details

Maturity
Pilot
Promoter
University of Thessaly (Dept. of Informatics & Telecommunications / Dept. of Special Education)
Period
2025–2026
Region (NUTS)
EL61
Keywords
special education, dyslexia screening, machine learning, EdTech

Context

Researchers at the University of Thessaly (Departments of Informatics & Telecommunications and Special Education) developed a deep-learning screening tool that analyses 7-second audio clips of Greek primary-school pupils reading aloud to flag potential reading difficulty.

Objectives

The tool is designed as a low-burden screening aid for teachers — not a diagnostic instrument — to help identify pupils who may need further evaluation for reading difficulties or dyslexia.

Activities

Teachers record pupils reading aloud; the audio is segmented into 7-second clips, converted to spectrograms, and classified by a YOLO-based convolutional neural network, producing an aggregate "difficulty percentage" score for each pupil.

Results

In a validation sample of 77 pupils (grades 3-6, mean age 11.6, including 6 with dyslexia diagnoses), the tool achieved about 84% overall accuracy, 0.85 balanced accuracy, and 0.89 macro F1, with a 0.74 correlation to expert judgment, using an 80/20 participant-level data split.

Conclusions

The authors state the tool needs further validation — including classroom-noise robustness testing, larger and more diverse samples, and multilingual testing — before wider implementation, and stress it should function only as a screening aid rather than a diagnostic tool.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful