Singapore's Department of Statistics (DOS) operates the SingStat Table Builder, a repository of roughly 2,400 data tables drawn from 70 public-sector agencies covering the country's economy and society. In 2025, DOS worked with design consultancy PebbleRoad to build SANDRA (Statistics ANd Data Retrieval Assistant), a conversational AI chatbot intended to close the gap between how the public actually searches — in plain language and concepts — and how the underlying datasets are technically named and organised; the project's own discovery work, based on search-log analysis and user interviews, documented that users habitually phrased queries in natural language rather than technical terms.
To build SANDRA, the team converted metadata from around 1,000 of the Table Builder's time-series datasets into vector embeddings, working with domain experts to preserve statistical meaning, and deployed the resulting semantic-search-plus-large-language-model system on AWS. Users can now ask questions in everyday English and receive interactive charts and tables in response, with the system also suggesting related datasets. DOS staff, per the case study, have been able to redirect time from answering basic look-up queries toward more complex requests.
SANDRA launched as a public beta — the PebbleRoad case study documenting the build is dated 21 August 2025, and the tool remains live at singstat.gov.sg/chatwithsandra-beta — but neither DOS nor PebbleRoad has published usage, adoption or accuracy figures since launch, and only a subset (about 1,000 of 2,400) of Table Builder's datasets have been vectorised so far. It is a well-documented example of a design-led approach to AI-based open-data discovery, though — like India's MCP pilot — one still awaiting a published post-launch evaluation.
Read the full analysis: https://www.singstat.gov.sg/chatwithsandra-beta
Where this practice's information was retrieved from, and when.