iFLYTEK AI Speech Scoring for China's School Oral English Exams
China
iFLYTEK's speech-recognition AI scores spoken English in China's zhongkao and gaokao oral exams across dozens of provinces, processing millions of …
Philippines · San Fernando (Pampanga) · See the Philippines profile
Evidence: Quasi-experimental Top 73% 40/100 · Ask Evidence Copilot about this practice
RTI's 2022 USAID/DepEd pilot tested AI speech-recognition scoring of oral reading fluency in 42 Philippine schools. The AI undercounted words-per-minute by 21.8 (English) / 25.1 (Filipino) vs. human graders, and DepEd cancelled the planned scale-up.
CoBRA (Computer-Based Reading Assessment) was a pilot tool for USAID's All Children Reading program in the Philippines, combining a voice-recording plugin, Google's Speech Recognition Engine for Filipino, and a machine-learning scoring algorithm to calculate words-correct-per-minute and reading accuracy from students' recorded oral reading, benchmarked against the Phil-IRI standard. RTI International developed it February-April 2022 and piloted it with DepEd June-July 2022 across 42 schools.
Test whether AI speech-recognition scoring could substitute for human assessment of oral reading fluency in English and Filipino at national scale.
The pilot tested 1,063 learners in English and 933 in Filipino across Grades 4-6, and trained or oriented 145 teachers and 67 ICT coordinators. RTI's final report (submitted November 2022) ran a concurrent-validity analysis, independently re-grading a subsample of 346 English and 345 Filipino assessments by hand and comparing them to the AI's scores.
The AI undercounted correct words-per-minute by an average of 21.8 in English and 25.1 in Filipino relative to human regraders, and incorrectly flagged or skipped 21% of English text and 60% of Filipino text that students actually read. Grade-level score variance between AI and human scoring ranged from 31-48% in English and 37-64% in Filipino, worsening as passages grew more complex. Over 90% of the 81 teachers surveyed found the tool easy to administer and felt the scoring was 'generally accurate' -- a perception the quantitative validity data did not support.
DepEd's planned larger second-phase pilot was cancelled, officially attributed to institutional reorganisation and unsustainable licensing costs; RTI's own report concluded that AI scoring 'is not accurate or reliable enough to use as an independent means for measuring students' oral reading fluency skills.' The practice is catalogued as a rigorously evaluated, honestly reported negative finding.
Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.
Where this practice's information was retrieved from, and when.
China
iFLYTEK's speech-recognition AI scores spoken English in China's zhongkao and gaokao oral exams across dozens of provinces, processing millions of …
United States of America
New Mexico mandates Amira's AI voice-based reading test for all K-2 students, but after parents raised biometric-privacy concerns, six districts …
Australia
Instead of an unwinnable AI-detection arms race, Sydney redesigned assessment into two lanes — secure in-person assessment of core capability, …
Croatia
CARNET's ESF+-funded BrAIn project piloted Croatia's first AI curriculum in VET schools, covering responsible AI use, plagiarism and authorship crediting; …
Open full copilot Grounded in cited practices — always check the sources.