evidoria

← Back to browse

Good practice Imported

iFLYTEK AI Speech Scoring for China's School Oral English Exams

China · Hefei · See the China profile

Evidence: Observational / pre–post Top 59% 47/100 · Ask Evidence Copilot about this practice

iFLYTEK's speech-recognition AI scores spoken English in China's zhongkao and gaokao oral exams across dozens of provinces, processing millions of students a year. Independent Rasch-model research supports scoring reliability, but transparency, appeals process and equity impacts

~770,000
Students tested per round (Jiangsu zhongkao oral English)
3,200+ venues across 13 cities
Exam venues per round (Jiangsu)
29
Provinces/cities covered (gaokao oral component)
ICC 0.74-0.92, r 0.85-0.87
Agreement with trained human raters (2 of 3 benchmarked tools) (2025)
2009
Deployment start (Jiangsu pilot)
iFLYTEK AI Speech Scoring for China's School Oral English Exams

Details

Maturity
Established
Promoter
iFLYTEK Co., Ltd. (科大讯飞)
Period
2009-present
Keywords
speech recognition, automated scoring, oral English assessment, exam technology

Context

iFLYTEK (科大讯飞), a Hefei-based speech-AI company, supplies the automated 'human-machine dialogue' (人机对话) scoring engine used in oral English components of China's zhongkao (senior-high entrance exam) and gaokao (national college entrance exam) in dozens of provinces.

Objectives

The system aims to grade pronunciation accuracy, fluency, completeness and content of spoken English answers in real time, replacing or supplementing panels of human examiners at a scale no manual process could match.

Activities

Students read aloud, answer questions or describe a picture into a headset; iFLYTEK's automatic speech recognition and scoring models grade the response. Jiangsu piloted computer-delivered oral English in its zhongkao from around 2009 and now tests roughly 770,000 middle-school students across 3,200+ venues in 13 cities each round; Xiamen and Tongling (Anhui) run comparable exams; iFLYTEK states its systems now cover the gaokao oral component in 29 provinces and cities.

Results

China's Ministry of Education has recognised iFLYTEK's system as the only application judged feasible and trustworthy for organising large-scale online spoken-language exams. Independent academic evidence on scoring validity is mixed but real: a 2016 multifaceted Rasch-model study found the automatic scoring reliable for junior-secondary zhongkao oral tests, while an earlier 2010 study found some students achieving high AI-assigned scores despite actual speaking proficiency below the required standard. A 2025 PLoS ONE study benchmarking Chinese automated spoken-English scoring tools against trained human raters (n=30 students) found two of three tools achieved strong agreement (ICC 0.74-0.92, r 0.85-0.87) while one showed systematic score inflation.

Conclusions

Transparency remains the weakest point in the public record: no independent audit of accent or dialect bias was found, the scoring algorithm and its weighting are proprietary, and no formal student-appeals process for AI-scored oral exams is documented. Company claims such as processing 20 million test-takers with zero complaints are self-reported marketing rather than independently audited figures.

Implementation

Indicative cost
High (€500k–€5M) — No published cost figures found; sustained deployment across 3,200+ venues in Jiangsu alone implies substantial vendor and provincial-government investment in equipment and licensing since 2009.
Time to results
Long (> 3 years) — Piloted in Jiangsu from around 2009; now covers the gaokao oral component in 29 provinces and cities, run continuously each exam cycle.
Staffing & skills
iFLYTEK speech-recognition/AI engineering teams, provincial education examination authorities administering exam venues, human examiners being replaced or supplemented for oral assessment

Conditions for success

  • reliable ASR infrastructure and headset equipment at scale across thousands of venues
  • Ministry of Education endorsement and regulatory acceptance
  • standardised exam format (read-aloud, Q&A, picture description) suited to automated grading

Common failure modes

  • proprietary, non-transparent scoring algorithm and weighting
  • no documented independent audit of accent or dialect bias
  • no formal student-appeals process for AI-scored oral exams
  • a 2010 study and one 2025-benchmarked tool both showed scoring inflation relative to true proficiency

Where it fits

Governance type
vendor-supplied exam infrastructure under Ministry of Education endorsement
Scale
national, dozens of provinces, hundreds of thousands of students per exam round
Income level
upper-middle-income (China)

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful