evidoria

← Back to browse

Good practice Imported

KemenkeuGPT — A Retrieval-Augmented LLM Pilot for Indonesia's Ministry of Finance

Indonesia · Jakarta · See the Indonesia profile · See the Jakarta profile

Evidence: Observational / pre–post Top 66% 53/100 · Ask Evidence Copilot about this practice

Researchers with Indonesia's Ministry of Finance built KemenkeuGPT, a retrieval-augmented LLM trained on 2003–2023 financial data from the Ministry, BPS and the IMF; fine-tuning raised accuracy from 35% to 61%.

35% → 61%
Answer accuracy (fine-tuned vs. baseline)
48% → 64%
Answer correctness (fine-tuned vs. baseline)
44%
RAGAS correctness score
73%
RAGAS faithfulness score
40%
RAGAS precision score
60%
RAGAS recall score

Details

Maturity
Pilot
Promoter
Ministry of Finance of the Republic of Indonesia (Kementerian Keuangan), developed with academic researchers
Period
2023–2024
Keywords
public financial management, budget policy, decision support

Context

Officials at Indonesia's Ministry of Finance face the recurring challenge of navigating two decades of complex, fragmented financial data and regulations to support evidence-based budget and policy decisions.

Objectives

Researchers working with the Ministry aimed to build a tool that could help officials query and reason over this fragmented multi-source corpus to support decision-making.

Activities

Researchers built KemenkeuGPT, a large language model system applying retrieval-augmented generation (RAG) via LangChain, prompt engineering and fine-tuning on a corpus spanning 2003–2023, drawn from the Ministry of Finance itself, Statistics Indonesia (BPS) and the IMF. Surveys and interviews with Ministry officials were used to inform and refine the model's outputs.

Results

Fine-tuning raised KemenkeuGPT's answer accuracy from 35% to 61% and correctness from 48% to 64%; under the RAGAS evaluation framework it reached 44% correctness, 73% faithfulness, 40% precision and 60% recall, outperforming several base models tested. An interviewed Ministry official described the tool as having potential to become 'an essential tool for decision-making'.

Conclusions

The published research frames KemenkeuGPT as a pilot and potential future tool rather than a system already in ministry-wide production use. Separately, in December 2024 the Ministry's Deputy Minister publicly announced plans to expand AI use across the institution — including anomaly detection in financial reports and correlation analysis between budget allocation and outcomes — signalling institutional appetite for this kind of tool beyond the research pilot itself.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year) — Research and development spanned 2023–2024; in December 2024 the Ministry's Deputy Minister cited it as grounds for wider institutional AI-expansion plans.
Staffing & skills
Academic researchers (RAG system development via LangChain, fine-tuning), Ministry of Finance officials (surveyed and interviewed to refine model outputs)

Conditions for success

  • Access to a structured multi-year financial corpus (Ministry of Finance, BPS, IMF, 2003–2023)
  • Combining fine-tuning with retrieval-augmented generation to raise accuracy over baseline models
  • Ministry official engagement via surveys/interviews to validate outputs

Common failure modes

  • Framed by its own authors as a pilot, with no confirmed ministry-wide production deployment
  • No documented data-protection, oversight or accountability framework for the system

Commonly funded by

National / regional programmes

Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.

Do you run this practice? Claim it — verified implementers get a public contact pathway and can propose corrections.

Data sources

Where this practice's information was retrieved from, and when.

Similar practices you may find useful