Study Finds ChatGPT Far Inferior to Traditional Models at Forecasting Inflation
Researchers from the Federal Reserve Bank of San Francisco and universities used ChatGPT-4 Turbo to forecast U.S. inflation and compared its performance with traditional econometric models. Inflation forecasts shape Federal Reserve interest-rate decisions and fund asset allocation. If AI performs well only in backtests where the answers are already known but cannot handle qualitative factors and random shocks, it is unlikely to replace the judgment of economists and portfolio managers.
The paper’s latest version was dated January 27, 2026, and posted on EERN on February 3. The results showed that ChatGPT-4 Turbo’s out-of-sample forecast error was as much as 12 times that of traditional models, while its outputs often relied on outdated information. Although it approached the benchmark when tested with data known after the fact, those results were vulnerable to data leakage and hindsight bias.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.