evidoria

← Back to browse

Good practice

LLM Inflation Forecasting Pilot — the Czech National Bank's Public Test of AI Against Its Core Model

Czechia · Prague · See the Czechia profile

In a public blog experiment, the Czech National Bank pitted general-purpose LLMs (OpenAI o1, Grok 2) against its own core inflation model and market analysts: o1 scored the lowest overall forecast error (RMSE 5.28 vs the bank's 5.48), but the bank's model won decisively in 2023–2

5.28
RMSE, OpenAI o1 (full 2020Q1-2024Q3 sample)
5.48
RMSE, ČNB core model (full sample)
6.19
RMSE, Grok 2 (full sample)
6.63
RMSE, market analysts (full sample)
0.50
RMSE, ČNB core model (2023Q2-2024Q3 disinflation period)
1.81
RMSE, OpenAI o1 (2023Q2-2024Q3 disinflation period)
99.7%
Automated CPI category classification agreement (broadest basket level)
0.669
Quantile Regression Forest model RMSE at 3-month horizon
LLM Inflation Forecasting Pilot — the Czech National Bank's Public Test of AI Against Its Core Model

Details

Maturity
Pilot
Promoter
Czech National Bank (Česká národní banka, ČNB)
Period
Feb 2025 – Apr 2026
Keywords
central banking, monetary policy, inflation forecasting, large language models, generative AI

Context

Between February 2025 and April 2026, the Czech National Bank, under Governor Aleš Michl (a co-author of the published analysis), ran a public test of whether general-purpose large language models could forecast Czech consumer-price inflation one year ahead by simple prompting, benchmarking OpenAI's o1 and xAI's Grok 2 against the bank's own core forecasting model and the median forecast of financial-market analysts.

Activities

The comparison ran over the 2020 Q1-2024 Q3 sample. In a companion experiment, ČNB researchers used OpenAI text-embedding models plus ChatGPT-4o-mini to auto-classify online retailer prices into CPI basket categories for real-time inflation nowcasting. Separately, an internationally refereed working paper (WP 09/2026) built a Quantile Regression Forest model with a novel time-varying-weight scheme.

Results

Over the full sample, OpenAI's o1 produced the lowest overall root-mean-square error (RMSE 5.28), narrowly ahead of the ČNB's core model (5.48), and both clearly ahead of Grok 2 (6.19) and market analysts (6.63). The ranking flipped by sub-period: during the 2021-2023 inflation surge o1 (7.94) still edged out the core model (8.41), but in the 2023 Q2-2024 Q3 disinflation back to target the core model was far more accurate (RMSE 0.50) than o1 (1.81), Grok 2 (3.45) or analysts (2.31). Automated CPI classification agreed with the bank's manual assignment 99.7% of the time at the broadest basket level, falling to 89.2% and 80.1% at finer levels of detail. The refereed Quantile Regression Forest model beat random walk, AR, ARIMA and ensemble benchmarks at a 3-month horizon (RMSE 0.669 vs 0.896-0.739), a statistically significant improvement confirmed by Diebold-Mariano tests.

Conclusions

The bank itself frames this as 'a test and demonstration' and 'a support tool for analysts,' not a replacement for its core model, explicitly flagging LLM hallucination risk and the non-reproducibility of prompted outputs; independent trade-press coverage echoed these 'black-box' concerns. No production forecasting process has been changed as a result.

Implementation

Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
Czech National Bank research economists (including Governor Aleš Michl as co-author)

Conditions for success

  • Rigorous backtesting against the bank's own core model and market analysts across multiple sub-periods
  • Transparent publication of methodology, comparison tables and explicit limitations

Common failure modes

  • LLM hallucination risk and non-reproducibility of prompted outputs flagged by the bank itself
  • Core model significantly outperformed LLMs during the 2023-24 disinflation period
  • Independent trade press raised 'black-box' concerns

Where it fits

Governance type
central bank
Scale
national research pilot
Income level
high-income

Data sources

Where this practice's information was retrieved from, and when.

Attachments

Similar practices you may find useful