Large Language Models: An Applied Econometric Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ludwig, Jens, Mullainathan, Sendhil, Rambachan, Ashesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917126163398656
author Ludwig, Jens
Mullainathan, Sendhil
Rambachan, Ashesh
author_facet Ludwig, Jens
Mullainathan, Sendhil
Rambachan, Ashesh
contents Large language models (LLMs) enable researchers to analyze text at unprecedented scale and minimal cost. Researchers can now revisit old questions and tackle novel ones with rich data. We provide an econometric framework for realizing this potential in two empirical uses. For prediction problems -- forecasting outcomes from text -- valid conclusions require ``no training leakage'' between the LLM's training data and the researcher's sample, which can be enforced through careful model choice and research design. For estimation problems -- automating the measurement of economic concepts for downstream analysis -- valid downstream inference requires combining LLM outputs with a small validation sample to deliver consistent and precise estimates. Absent a validation sample, researchers cannot assess possible errors in LLM outputs, and consequently seemingly innocuous choices (which model, which prompt) can produce dramatically different parameter estimates. When used appropriately, LLMs are powerful tools that can expand the frontier of empirical economics.
format Preprint
id arxiv_https___arxiv_org_abs_2412_07031
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Models: An Applied Econometric Framework
Ludwig, Jens
Mullainathan, Sendhil
Rambachan, Ashesh
Econometrics
Artificial Intelligence
Large language models (LLMs) enable researchers to analyze text at unprecedented scale and minimal cost. Researchers can now revisit old questions and tackle novel ones with rich data. We provide an econometric framework for realizing this potential in two empirical uses. For prediction problems -- forecasting outcomes from text -- valid conclusions require ``no training leakage'' between the LLM's training data and the researcher's sample, which can be enforced through careful model choice and research design. For estimation problems -- automating the measurement of economic concepts for downstream analysis -- valid downstream inference requires combining LLM outputs with a small validation sample to deliver consistent and precise estimates. Absent a validation sample, researchers cannot assess possible errors in LLM outputs, and consequently seemingly innocuous choices (which model, which prompt) can produce dramatically different parameter estimates. When used appropriately, LLMs are powerful tools that can expand the frontier of empirical economics.
title Large Language Models: An Applied Econometric Framework
topic Econometrics
Artificial Intelligence
url https://arxiv.org/abs/2412.07031