EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910025908224000 |
|---|---|
| author | Fish, Sara Shephard, Julia Li, Minkai Shorrer, Ran I. Gonczarowski, Yannai A. |
| author_facet | Fish, Sara Shephard, Julia Li, Minkai Shorrer, Ran I. Gonczarowski, Yannai A. |
| contents | We develop evaluation methods for measuring the economic decision-making capabilities and tendencies of LLMs. First, we develop benchmarks derived from key problems in economics -- procurement, scheduling, and pricing -- that test an LLM's ability to learn from the environment in context. Second, we develop the framework of litmus tests, evaluations that quantify an LLM's choice behavior on a stylized decision-making task with multiple conflicting objectives. Each litmus test outputs a litmus score, which quantifies an LLM's tradeoff response, a reliability score, which measures the coherence of an LLM's choice behavior, and a competency score, which measures an LLM's capability at the same task when the conflicting objectives are replaced by a single, well-specified objective. Evaluating a broad array of frontier LLMs, we (1) investigate changes in LLM capabilities and tendencies over time, (2) derive economically meaningful insights from the LLMs' choice behavior and chain-of-thought, (3) validate our litmus test framework by testing self-consistency, robustness, and generalizability. Overall, this work provides a foundation for evaluating LLM agents as they are further integrated into economic decision-making. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_18825 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents Fish, Sara Shephard, Julia Li, Minkai Shorrer, Ran I. Gonczarowski, Yannai A. Artificial Intelligence Computation and Language Computer Science and Game Theory We develop evaluation methods for measuring the economic decision-making capabilities and tendencies of LLMs. First, we develop benchmarks derived from key problems in economics -- procurement, scheduling, and pricing -- that test an LLM's ability to learn from the environment in context. Second, we develop the framework of litmus tests, evaluations that quantify an LLM's choice behavior on a stylized decision-making task with multiple conflicting objectives. Each litmus test outputs a litmus score, which quantifies an LLM's tradeoff response, a reliability score, which measures the coherence of an LLM's choice behavior, and a competency score, which measures an LLM's capability at the same task when the conflicting objectives are replaced by a single, well-specified objective. Evaluating a broad array of frontier LLMs, we (1) investigate changes in LLM capabilities and tendencies over time, (2) derive economically meaningful insights from the LLMs' choice behavior and chain-of-thought, (3) validate our litmus test framework by testing self-consistency, robustness, and generalizability. Overall, this work provides a foundation for evaluating LLM agents as they are further integrated into economic decision-making. |
| title | EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents |
| topic | Artificial Intelligence Computation and Language Computer Science and Game Theory |
| url | https://arxiv.org/abs/2503.18825 |