STEER: Assessing the Economic Rationality of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929361596186624 |
|---|---|
| author | Raman, Narun Lundy, Taylor Amouyal, Samuel Levine, Yoav Leyton-Brown, Kevin Tennenholtz, Moshe |
| author_facet | Raman, Narun Lundy, Taylor Amouyal, Samuel Levine, Yoav Leyton-Brown, Kevin Tennenholtz, Moshe |
| contents | There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions -- and more broadly, determining whether an LLM agent is reliable enough to be trusted -- requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "STEER report card." Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_09552 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | STEER: Assessing the Economic Rationality of Large Language Models Raman, Narun Lundy, Taylor Amouyal, Samuel Levine, Yoav Leyton-Brown, Kevin Tennenholtz, Moshe Computation and Language General Economics Economics There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions -- and more broadly, determining whether an LLM agent is reliable enough to be trusted -- requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of fine-grained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "STEER report card." Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior. |
| title | STEER: Assessing the Economic Rationality of Large Language Models |
| topic | Computation and Language General Economics Economics |
| url | https://arxiv.org/abs/2402.09552 |