A peer-reviewed study applied topic modelling and a Random Forest classifier to 21,833 real proposals from Seoul's participatory budget (2013–2021), reaching 0.73 AUROC in predicting selected proposals — a research-stage tool to cut proposal-screening workload.
50 KRW billion
Annual participatory budget allocated by Seoul (annual, since 2012)
2.2 %
Share of city budget allocated via participatory budgeting (annual)
1,533 proposals
Proposals received in 2014 (2014)
352 proposals
Proposals selected for funding in 2014 (2014)
21,833 proposals
De-duplicated proposals analysed in the machine-learning study (2013–2021)
0.73 AUROC
Random Forest classifier accuracy (AUROC) predicting proposal selection (2026 study)
Details
Maturity
Pilot
Promoter
Seoul Metropolitan Government
Period
2013–2021 (PB data studied); ML study published 2026
Keywords
participatory budgeting, civic tech, public administration, machine learning research
Context
Seoul has run one of the world's largest participatory budgeting (PB) programmes since May 2012, allocating roughly 50 billion Korean won (about 2.2% of the city budget) through a General Committee of 250 citizens, officials and civil-society representatives; in 2014 alone the process received 1,533 proposals, of which 352 were ultimately selected for funding.
Objectives
To help handle this volume, researcher Bokyong Shin (University of Helsinki) published a 2026 study aiming to cut citizens' 'learning costs' through automated categorisation and summarisation, and 'compliance costs' through predictive feedback before formal deliberation.
Activities
The study analysed 21,833 de-duplicated proposals (from 25,415 originally submitted) spanning 2013–2021, combining unsupervised topic modelling using Sentence-BERT embeddings — identifying 17 major topics across two clusters (nine infrastructure-focused, eight social/welfare-focused) — with a supervised Random Forest classifier that predicted proposal selection.
Results
The Random Forest classifier predicted proposal selection with an AUROC of 0.73, outperforming logistic regression, gradient boosting, SVM and decision-tree alternatives.
Conclusions
This remains a research-stage capability built on historical open data rather than a tool confirmed to be running inside Seoul's live PB workflow, so its screening role should be read as a validated proof of concept, not an operational system. The programme's underlying scale and governance structure are independently corroborated by Participedia's case documentation.
Implementation
Indicative cost
Low (< €50k) — Research conducted on existing open administrative data; no operational deployment budget is reported in the source.
Time to results
Medium (1–3 years) — PB programme running since May 2012; proposal data studied spans 2013–2021; machine-learning study published 2026.
Staffing & skills
Seoul Metropolitan Government (PB programme owner), Academic researcher: Bokyong Shin, University of Helsinki
Conditions for success
Availability of historical PB proposal data with recorded selection outcomes for model training
A well-established, high-volume PB programme generating enough data to model (21,833 proposals over 2013–2021)
Common failure modes
Remains research-stage; not confirmed to be running inside Seoul's live PB workflow
An AUROC of 0.73 leaves substantial misclassification risk if deployed operationally without further validation
Where it fits
Governance type
city government with a citizen general committee
Scale
city-wide (Seoul)
Income level
high-income
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
Under a Scottish Government–COSLA framework requiring at least 1% of council budgets to go through participatory budgeting, researchers evaluated NLP …
Warwick, QMUL and Alan Turing Institute researchers built four NLP/ML tools — recommendation, interest grouping, summarisation, aggregation — for Madrid's …