Revealing economic facts: LLMs know more than they say

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Buckmann, Marcus, Nguyen, Quynh Anh, Hill, Edward
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917136130113536
author Buckmann, Marcus
Nguyen, Quynh Anh
Hill, Edward
author_facet Buckmann, Marcus
Nguyen, Quynh Anh
Hill, Edward
contents We investigate whether the hidden states of large language models (LLMs) can be used to estimate and impute economic and financial statistics. Focusing on county-level (e.g. unemployment) and firm-level (e.g. total assets) variables, we show that a simple linear model trained on the hidden states of open-source LLMs outperforms the models' text outputs. This suggests that hidden states capture richer economic information than the responses of the LLMs reveal directly. A learning curve analysis indicates that only a few dozen labelled examples are sufficient for training. We also propose a transfer learning method that improves estimation accuracy without requiring any labelled data for the target variable. Finally, we demonstrate the practical utility of hidden-state representations in super-resolution and data imputation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08662
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revealing economic facts: LLMs know more than they say
Buckmann, Marcus
Nguyen, Quynh Anh
Hill, Edward
Computation and Language
Machine Learning
General Economics
Economics
I.2.7
We investigate whether the hidden states of large language models (LLMs) can be used to estimate and impute economic and financial statistics. Focusing on county-level (e.g. unemployment) and firm-level (e.g. total assets) variables, we show that a simple linear model trained on the hidden states of open-source LLMs outperforms the models' text outputs. This suggests that hidden states capture richer economic information than the responses of the LLMs reveal directly. A learning curve analysis indicates that only a few dozen labelled examples are sufficient for training. We also propose a transfer learning method that improves estimation accuracy without requiring any labelled data for the target variable. Finally, we demonstrate the practical utility of hidden-state representations in super-resolution and data imputation tasks.
title Revealing economic facts: LLMs know more than they say
topic Computation and Language
Machine Learning
General Economics
Economics
I.2.7
url https://arxiv.org/abs/2505.08662