evidoria

← Back to browse

Good practice Imported

CoBRA — RTI International's AI Speech-Recognition Pilot for Oral Reading Assessment in the Philippines

Philippines · San Fernando (Pampanga) · See the Philippines profile

Evidence: Quasi-experimental Top 73% 40/100 · Ask Evidence Copilot about this practice

RTI's 2022 USAID/DepEd pilot tested AI speech-recognition scoring of oral reading fluency in 42 Philippine schools. The AI undercounted words-per-minute by 21.8 (English) / 25.1 (Filipino) vs. human graders, and DepEd cancelled the planned scale-up.

1,063 Grades 4-6
Learners tested, English (Jun-Jul 2022)
933 Grades 4-6
Learners tested, Filipino (Jun-Jul 2022)
21.8 avg words, vs human regraders
AI undercount of words-correct-per-minute, English
60%
Text incorrectly flagged/skipped, Filipino
CoBRA — RTI International's AI Speech-Recognition Pilot for Oral Reading Assessment in the Philippines

Details

Maturity
Discontinued
Promoter
RTI International; USAID All Children Reading (ACR)–Philippines; Philippines Department of Education (DepEd)
Period
February–November 2022 (development Feb–Apr; pilot Jun–Jul; final report Nov 2022)
Keywords
primary education, reading assessment, speech recognition, EdTech evaluation

Context

CoBRA (Computer-Based Reading Assessment) was a pilot tool for USAID's All Children Reading program in the Philippines, combining a voice-recording plugin, Google's Speech Recognition Engine for Filipino, and a machine-learning scoring algorithm to calculate words-correct-per-minute and reading accuracy from students' recorded oral reading, benchmarked against the Phil-IRI standard. RTI International developed it February-April 2022 and piloted it with DepEd June-July 2022 across 42 schools.

Objectives

Test whether AI speech-recognition scoring could substitute for human assessment of oral reading fluency in English and Filipino at national scale.

Activities

The pilot tested 1,063 learners in English and 933 in Filipino across Grades 4-6, and trained or oriented 145 teachers and 67 ICT coordinators. RTI's final report (submitted November 2022) ran a concurrent-validity analysis, independently re-grading a subsample of 346 English and 345 Filipino assessments by hand and comparing them to the AI's scores.

Results

The AI undercounted correct words-per-minute by an average of 21.8 in English and 25.1 in Filipino relative to human regraders, and incorrectly flagged or skipped 21% of English text and 60% of Filipino text that students actually read. Grade-level score variance between AI and human scoring ranged from 31-48% in English and 37-64% in Filipino, worsening as passages grew more complex. Over 90% of the 81 teachers surveyed found the tool easy to administer and felt the scoring was 'generally accurate' -- a perception the quantitative validity data did not support.

Conclusions

DepEd's planned larger second-phase pilot was cancelled, officially attributed to institutional reorganisation and unsustainable licensing costs; RTI's own report concluded that AI scoring 'is not accurate or reliable enough to use as an independent means for measuring students' oral reading fluency skills.' The practice is catalogued as a rigorously evaluated, honestly reported negative finding.

Implementation

Indicative cost
Medium (€50k–€500k)
Time to results
Short (< 1 year)
Staffing & skills
RTI International development and evaluation team, 145 trained teachers and 67 ICT coordinators (DepEd)

Conditions for success

  • Benchmarking against an established national reading-assessment standard (Phil-IRI)
  • Independent human re-grading of a representative subsample for validity testing

Common failure modes

  • Substantial undercounting and misflagging of text vs human graders
  • Teacher perception of accuracy not supported by the quantitative validity data
  • Second-phase pilot cancelled due to reorganisation and unsustainable licensing costs

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful