Researchers with Indonesia's Ministry of Finance built KemenkeuGPT, a retrieval-augmented LLM trained on 2003–2023 financial data from the Ministry, BPS and the IMF; fine-tuning raised accuracy from 35% to 61%.
35% → 61%
Answer accuracy (fine-tuned vs. baseline)
48% → 64%
Answer correctness (fine-tuned vs. baseline)
44%
RAGAS correctness score
73%
RAGAS faithfulness score
40%
RAGAS precision score
60%
RAGAS recall score
Details
Maturity
Pilot
Promoter
Ministry of Finance of the Republic of Indonesia (Kementerian Keuangan), developed with academic researchers
Period
2023–2024
Keywords
public financial management, budget policy, decision support
Context
Officials at Indonesia's Ministry of Finance face the recurring challenge of navigating two decades of complex, fragmented financial data and regulations to support evidence-based budget and policy decisions.
Objectives
Researchers working with the Ministry aimed to build a tool that could help officials query and reason over this fragmented multi-source corpus to support decision-making.
Activities
Researchers built KemenkeuGPT, a large language model system applying retrieval-augmented generation (RAG) via LangChain, prompt engineering and fine-tuning on a corpus spanning 2003–2023, drawn from the Ministry of Finance itself, Statistics Indonesia (BPS) and the IMF. Surveys and interviews with Ministry officials were used to inform and refine the model's outputs.
Results
Fine-tuning raised KemenkeuGPT's answer accuracy from 35% to 61% and correctness from 48% to 64%; under the RAGAS evaluation framework it reached 44% correctness, 73% faithfulness, 40% precision and 60% recall, outperforming several base models tested. An interviewed Ministry official described the tool as having potential to become 'an essential tool for decision-making'.
Conclusions
The published research frames KemenkeuGPT as a pilot and potential future tool rather than a system already in ministry-wide production use. Separately, in December 2024 the Ministry's Deputy Minister publicly announced plans to expand AI use across the institution — including anomaly detection in financial reports and correlation analysis between budget allocation and outcomes — signalling institutional appetite for this kind of tool beyond the research pilot itself.
Implementation
Indicative cost
Low (< €50k)
Time to results
Short (< 1 year) — Research and development spanned 2023–2024; in December 2024 the Ministry's Deputy Minister cited it as grounds for wider institutional AI-expansion plans.
Staffing & skills
Academic researchers (RAG system development via LangChain, fine-tuning), Ministry of Finance officials (surveyed and interviewed to refine model outputs)
Conditions for success
Access to a structured multi-year financial corpus (Ministry of Finance, BPS, IMF, 2003–2023)
Combining fine-tuning with retrieval-augmented generation to raise accuracy over baseline models
Ministry official engagement via surveys/interviews to validate outputs
Common failure modes
Framed by its own authors as a pilot, with no confirmed ministry-wide production deployment
No documented data-protection, oversight or accountability framework for the system
Commonly funded by
National / regional programmes
Indicative funding routes for practices of this type — always check each programme's current calls and eligibility rules.
Do you run this practice?
Claim it —
verified implementers get a public contact pathway and can propose corrections.
Data sources
Where this practice's information was retrieved from, and when.
Bangladesh's a2i, GiveDirectly, UNDP and UC Berkeley piloted ML poverty-targeting from call-detail records across 106,200 Cox's Bazar households, but a …