iFLYTEK's speech-recognition AI scores spoken English in China's zhongkao and gaokao oral exams across dozens of provinces, processing millions of students a year. Independent Rasch-model research supports scoring reliability, but transparency, appeals process and equity impacts
~770,000
Students tested per round (Jiangsu zhongkao oral English)
3,200+ venues across 13 cities
Exam venues per round (Jiangsu)
29
Provinces/cities covered (gaokao oral component)
ICC 0.74-0.92, r 0.85-0.87
Agreement with trained human raters (2 of 3 benchmarked tools) (2025)
2009
Deployment start (Jiangsu pilot)
Details
Maturity
Established
Promoter
iFLYTEK Co., Ltd. (科大讯飞)
Period
2009-present
Keywords
speech recognition, automated scoring, oral English assessment, exam technology
Context
iFLYTEK (科大讯飞), a Hefei-based speech-AI company, supplies the automated 'human-machine dialogue' (人机对话) scoring engine used in oral English components of China's zhongkao (senior-high entrance exam) and gaokao (national college entrance exam) in dozens of provinces.
Objectives
The system aims to grade pronunciation accuracy, fluency, completeness and content of spoken English answers in real time, replacing or supplementing panels of human examiners at a scale no manual process could match.
Activities
Students read aloud, answer questions or describe a picture into a headset; iFLYTEK's automatic speech recognition and scoring models grade the response. Jiangsu piloted computer-delivered oral English in its zhongkao from around 2009 and now tests roughly 770,000 middle-school students across 3,200+ venues in 13 cities each round; Xiamen and Tongling (Anhui) run comparable exams; iFLYTEK states its systems now cover the gaokao oral component in 29 provinces and cities.
Results
China's Ministry of Education has recognised iFLYTEK's system as the only application judged feasible and trustworthy for organising large-scale online spoken-language exams. Independent academic evidence on scoring validity is mixed but real: a 2016 multifaceted Rasch-model study found the automatic scoring reliable for junior-secondary zhongkao oral tests, while an earlier 2010 study found some students achieving high AI-assigned scores despite actual speaking proficiency below the required standard. A 2025 PLoS ONE study benchmarking Chinese automated spoken-English scoring tools against trained human raters (n=30 students) found two of three tools achieved strong agreement (ICC 0.74-0.92, r 0.85-0.87) while one showed systematic score inflation.
Conclusions
Transparency remains the weakest point in the public record: no independent audit of accent or dialect bias was found, the scoring algorithm and its weighting are proprietary, and no formal student-appeals process for AI-scored oral exams is documented. Company claims such as processing 20 million test-takers with zero complaints are self-reported marketing rather than independently audited figures.
Implementation
Indicative cost
High (€500k–€5M) — No published cost figures found; sustained deployment across 3,200+ venues in Jiangsu alone implies substantial vendor and provincial-government investment in equipment and licensing since 2009.
Time to results
Long (> 3 years) — Piloted in Jiangsu from around 2009; now covers the gaokao oral component in 29 provinces and cities, run continuously each exam cycle.
Staffing & skills
iFLYTEK speech-recognition/AI engineering teams, provincial education examination authorities administering exam venues, human examiners being replaced or supplemented for oral assessment
Conditions for success
reliable ASR infrastructure and headset equipment at scale across thousands of venues
Ministry of Education endorsement and regulatory acceptance
standardised exam format (read-aloud, Q&A, picture description) suited to automated grading
Common failure modes
proprietary, non-transparent scoring algorithm and weighting
no documented independent audit of accent or dialect bias
no formal student-appeals process for AI-scored oral exams
a 2010 study and one 2025-benchmarked tool both showed scoring inflation relative to true proficiency
Where it fits
Governance type
vendor-supplied exam infrastructure under Ministry of Education endorsement
Scale
national, dozens of provinces, hundreds of thousands of students per exam round
Income level
upper-middle-income (China)
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
RTI's 2022 USAID/DepEd pilot tested AI speech-recognition scoring of oral reading fluency in 42 Philippine schools. The AI undercounted words-per-minute …
CARNET's ESF+-funded BrAIn project piloted Croatia's first AI curriculum in VET schools, covering responsible AI use, plagiarism and authorship crediting; …