STEER: Assessing the Economic Rationality of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raman, Narun, Lundy, Taylor, Amouyal, Samuel, Levine, Yoav, Leyton-Brown, Kevin, Tennenholtz, Moshe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929361596186624
author Raman, Narun
Lundy, Taylor
Amouyal, Samuel
Levine, Yoav
Leyton-Brown, Kevin
Tennenholtz, Moshe
author_facet Raman, Narun
Lundy, Taylor
Amouyal, Samuel
Levine, Yoav
Leyton-Brown, Kevin
Tennenholtz, Moshe
contents There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions -- and more broadly, determining whether an LLM agent is reliable enough to be trusted -- requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "STEER report card." Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09552
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle STEER: Assessing the Economic Rationality of Large Language Models
Raman, Narun
Lundy, Taylor
Amouyal, Samuel
Levine, Yoav
Leyton-Brown, Kevin
Tennenholtz, Moshe
Computation and Language
General Economics
Economics
There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions -- and more broadly, determining whether an LLM agent is reliable enough to be trusted -- requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "STEER report card." Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior.
title STEER: Assessing the Economic Rationality of Large Language Models
topic Computation and Language
General Economics
Economics
url https://arxiv.org/abs/2402.09552