evidoria

← Back to browse

Good practice Imported

Seoul's Machine-Learning Pipeline for Screening Participatory Budgeting Proposals

South Korea · Seoul · See the South Korea profile

A peer-reviewed study applied topic modelling and a Random Forest classifier to 21,833 real proposals from Seoul's participatory budget (2013–2021), reaching 0.73 AUROC in predicting selected proposals — a research-stage tool to cut proposal-screening workload.

Details

Promoter
Seoul Metropolitan Government
Period
2013–2021 (PB data studied); ML study published 2026
Keywords
participatory budgeting, civic tech, public administration, machine learning research

Description

Seoul has run one of the world's largest participatory budgeting (PB) programmes since May 2012, allocating roughly 50 billion Korean won (about 2.2% of the city budget) through a General Committee of 250 citizens, officials and civil-society representatives; in 2014 alone the process received 1,533 proposals, of which 352 were ultimately selected for funding.

To help handle this volume, researcher Bokyong Shin (University of Helsinki) published a 2026 study in the Journal of Public Budgeting, Accounting & Financial Management analysing 21,833 de-duplicated proposals (from 25,415 originally submitted) spanning 2013–2021. The study combined unsupervised topic modelling using Sentence-BERT embeddings — identifying 17 major topics across two clusters (nine infrastructure-focused, eight social/welfare-focused) — with a supervised Random Forest classifier that predicted proposal selection with an AUROC of 0.73, outperforming logistic regression, gradient boosting, SVM and decision-tree alternatives. The aim is to cut citizens' 'learning costs' through automated categorisation and summarisation, and 'compliance costs' through predictive feedback before formal deliberation.

This remains a research-stage capability built on historical open data rather than a tool confirmed to be running inside Seoul's live PB workflow, so its screening role should be read as a validated proof of concept, not an operational system. The programme's underlying scale and governance structure are independently corroborated by Participedia's case documentation.

Read the full analysis: https://www.emerald.com/jpbafm/article/38/1/237/1300647/Exploring-the-potential-of-machine-learning-to

Implementation

Implementation detail (cost, timeline, staffing, conditions for success) is not yet available for this practice.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful