Singapore's Department of Statistics launched SANDRA, a semantic-search AI chatbot for querying its 2,400-table SingStat Table Builder in plain English. Built by vectorising about 1,000 datasets, it remains in public beta with no published usage figures yet.
2400 tables
Data tables in SingStat Table Builder
70 agencies
Public-sector agencies contributing data
1000 datasets
Datasets vectorised for SANDRA (as of 2025)
Details
Maturity
Pilot
Promoter
Singapore Department of Statistics (DOS), with design consultancy PebbleRoad
Period
Built and launched as a public beta in 2025 (case study published 21 August 2025); ongoing
Keywords
digital government, open data, official statistics, conversational AI
Context
Singapore's Department of Statistics (DOS) manages the SingStat Table Builder, a repository of roughly 2,400 data tables sourced from 70 public-sector agencies. Recognising that the public tends to search using everyday language rather than technical dataset names, DOS partnered with design consultancy PebbleRoad to close that gap.
Objectives
The goal was to let users query official statistics in plain English and receive relevant charts, tables and related-dataset suggestions, while freeing DOS staff from routine look-up queries.
Activities
The team converted metadata from about 1,000 of the Table Builder's time-series datasets into vector embeddings, working with domain experts to preserve statistical meaning, and deployed a semantic-search-plus-LLM chatbot called SANDRA on AWS.
Conclusions
SANDRA has been live as a public beta since August 2025, but DOS and PebbleRoad have not yet published usage, adoption or accuracy figures, and only a subset of datasets have been vectorised so far, so the initiative remains an unevaluated but well-documented design-led pilot.
Implementation
Indicative cost
Medium (€50k–€500k)
Time to results
Short (< 1 year)
Staffing & skills
Department of Statistics (DOS) staff, PebbleRoad design consultancy
Conditions for success
Search-log analysis and user interviews to ground design in real query behaviour
Domain experts involved in vectorising metadata to preserve statistical meaning
Cloud (AWS) hosting for the semantic-search-plus-LLM system
Common failure modes
Only a subset (about 1,000 of 2,400) of datasets vectorised so far, limiting coverage
No published usage, adoption or accuracy evaluation since launch
Where it fits
Governance type
national statistics agency
Scale
national
Income level
high
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
India's National Statistical Office launched a public Model Context Protocol server letting AI assistants query official economic and social statistics …