In a public blog experiment, the Czech National Bank pitted general-purpose LLMs (OpenAI o1, Grok 2) against its own core inflation model and market analysts: o1 scored the lowest overall forecast error (RMSE 5.28 vs the bank's 5.48), but the bank's model won decisively in 2023–2
5.28
RMSE, OpenAI o1 (full 2020Q1-2024Q3 sample)
5.48
RMSE, ČNB core model (full sample)
6.19
RMSE, Grok 2 (full sample)
6.63
RMSE, market analysts (full sample)
0.50
RMSE, ČNB core model (2023Q2-2024Q3 disinflation period)
Quantile Regression Forest model RMSE at 3-month horizon
Details
Maturity
Pilot
Promoter
Czech National Bank (Česká národní banka, ČNB)
Period
Feb 2025 – Apr 2026
Keywords
central banking, monetary policy, inflation forecasting, large language models, generative AI
Context
Between February 2025 and April 2026, the Czech National Bank, under Governor Aleš Michl (a co-author of the published analysis), ran a public test of whether general-purpose large language models could forecast Czech consumer-price inflation one year ahead by simple prompting, benchmarking OpenAI's o1 and xAI's Grok 2 against the bank's own core forecasting model and the median forecast of financial-market analysts.
Activities
The comparison ran over the 2020 Q1-2024 Q3 sample. In a companion experiment, ČNB researchers used OpenAI text-embedding models plus ChatGPT-4o-mini to auto-classify online retailer prices into CPI basket categories for real-time inflation nowcasting. Separately, an internationally refereed working paper (WP 09/2026) built a Quantile Regression Forest model with a novel time-varying-weight scheme.
Results
Over the full sample, OpenAI's o1 produced the lowest overall root-mean-square error (RMSE 5.28), narrowly ahead of the ČNB's core model (5.48), and both clearly ahead of Grok 2 (6.19) and market analysts (6.63). The ranking flipped by sub-period: during the 2021-2023 inflation surge o1 (7.94) still edged out the core model (8.41), but in the 2023 Q2-2024 Q3 disinflation back to target the core model was far more accurate (RMSE 0.50) than o1 (1.81), Grok 2 (3.45) or analysts (2.31). Automated CPI classification agreed with the bank's manual assignment 99.7% of the time at the broadest basket level, falling to 89.2% and 80.1% at finer levels of detail. The refereed Quantile Regression Forest model beat random walk, AR, ARIMA and ensemble benchmarks at a 3-month horizon (RMSE 0.669 vs 0.896-0.739), a statistically significant improvement confirmed by Diebold-Mariano tests.
Conclusions
The bank itself frames this as 'a test and demonstration' and 'a support tool for analysts,' not a replacement for its core model, explicitly flagging LLM hallucination risk and the non-reproducibility of prompted outputs; independent trade-press coverage echoed these 'black-box' concerns. No production forecasting process has been changed as a result.
Implementation
Indicative cost
Low (< €50k)
Time to results
Short (< 1 year)
Staffing & skills
Czech National Bank research economists (including Governor Aleš Michl as co-author)
Conditions for success
Rigorous backtesting against the bank's own core model and market analysts across multiple sub-periods
Transparent publication of methodology, comparison tables and explicit limitations
Common failure modes
LLM hallucination risk and non-reproducibility of prompted outputs flagged by the bank itself
Core model significantly outperformed LLMs during the 2023-24 disinflation period
SARB researchers rigorously tested machine-learning models (LASSO, XGBoost, neural networks) against the bank's own published inflation forecasts: ML modestly beat …
Bundesbank researchers built MILA, an LLM agent that classifies the tone of ECB Governing Council communications sentence-by-sentence from 2011–2024 against …