evidoria

← Back to browse

Good practice Imported

Gemeentelijke Chatbots Getest — an Independent Audit Finds Dutch Municipal AI Chatbots Answer Only 1 in 10 Questions Correctly

Netherlands · The Hague · See the Netherlands profile · See the The Hague profile

Evidence: Observational / pre–post Top 88% 33/100 · Ask Evidence Copilot about this practice

An independent second-year audit of 39 Dutch municipal chatbots found only 10% of citizen questions answered correctly, with 64% wrong and identical errors repeated across ~30 municipalities sharing one 'Gem' chatbot; VNG called the tools 'not yet finished products.'

39
Municipal chatbots tested (2nd edition, Dec 2025-Feb 2026)
8
Questions per chatbot
312
Total question-answer exchanges
64 %
Answers wrong
23 %
Answers returning only a hyperlink
10 %
Answers correct
3 %
No response produced
0 of 39
Chatbots correctly answering the waste-collection question
~30
Municipalities sharing the 'Gem' chatbot
at most 10 of 39
Chatbots estimated to be genuinely LLM-powered
Gemeentelijke Chatbots Getest — an Independent Audit Finds Dutch Municipal AI Chatbots Answer Only 1 in 10 Questions Correctly

Details

Maturity
Established
Promoter
Independent researchers Wiep Hamstra and Jules Ernst; response from VNG (Vereniging van Nederlandse Gemeenten) and the Ministry of the Interior and Kingdom Relations (BZK)
Period
Fieldwork December 2025 - early February 2026; published 9 March 2026 (second annual edition; first edition June 2025)
Keywords
chatbots, citizen services, digital accessibility, government oversight, generative AI evaluation

Context

Since June 2025, independent researchers Wiep Hamstra and Jules Ernst have run an unfunded audit of municipal chatbots deployed across Dutch local government, after early evidence suggested these AI-branded tools might not perform as advertised.

Objectives

To systematically and repeatably test how accurately municipal chatbots answer real citizen questions, and track whether performance improves over time.

Activities

For the second edition (fieldwork December 2025-early February 2026, published 9 March 2026), the researchers tested 39 municipal chatbots with the same eight citizen questions each - a mix of routine service queries (bulky-waste collection, permit status) and election-related questions (replacing a lost voting pass) - for 312 total question-answer exchanges.

Results

Results were largely unchanged from the first edition: 64% of answers were wrong, 23% returned only a hyperlink (often to a page without the answer), 10% were correct, and 3% produced no response. Not one of the 39 chatbots correctly answered a basic household waste-collection question unaided. The researchers estimate at most 10 of the 39 tools are actually powered by a language model; most rely on hand-maintained, scripted question-answer pairs. Because nearly 30 municipalities license the same ‘Gem’ chatbot product, identical wrong answers propagate across all of them simultaneously.

Conclusions

VNG, the association representing Dutch municipalities, publicly acknowledged the findings, stating the new AI chatbots “are still no finished products” and that data quality determines whether a bot gives an outdated or hallucinated answer; it also noted chatbots have reduced call-centre volume, a claim the audit did not independently test. As a recurring, methodologically consistent, independently run check on a widely deployed public-facing AI tool, the audit is a rare case of rigorous accountability for municipal AI - even though what it evidences is mostly failure, not success.

Implementation

Indicative cost
Low (< €50k)
Time to results
Medium (1–3 years) — First edition published June 2025; second edition fieldwork December 2025-early February 2026, published 9 March 2026.
Staffing & skills
Conducted by two independent researchers, Wiep Hamstra (communications) and Jules Ernst (digital accessibility), on an unfunded basis, VNG (association of Dutch municipalities) and the Ministry of the Interior and Kingdom Relations (BZK) responded publicly but did not run the audit themselves

Conditions for success

  • Standardised set of eight citizen questions applied identically across all 39 chatbots, repeated for a second consecutive year for comparability
  • Findings made public with named methodology, enabling accountability pressure on municipalities and the shared 'Gem' platform vendor

Common failure modes

  • 64% of the 312 tested answers were wrong, 23% returned only a (often unhelpful) hyperlink, 3% produced no response, and only 10% were correct
  • Not one of the 39 chatbots correctly answered a basic household waste-collection question unaided
  • At most 10 of 39 tools are estimated to actually be powered by a language model; most rely on hand-maintained scripted Q&A pairs, so the 'AI chatbot' framing overstates the underlying technology
  • Because nearly 30 municipalities license the same shared 'Gem' chatbot, identical wrong answers propagate across all of them simultaneously
  • VNG itself acknowledged the tools are 'still no finished products' and that data-quality gaps drive both outdated and hallucinated answers

Where it fits

Governance type
independent civil-society audit of municipal government tools
Scale
national - 39 municipalities, ~30 sharing one shared chatbot product
Income level
high-income

Replication kit

Reusable artefacts from this practice — as published by their sources.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful