EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fish, Sara, Shephard, Julia, Li, Minkai, Shorrer, Ran I., Gonczarowski, Yannai A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910025908224000
author Fish, Sara
Shephard, Julia
Li, Minkai
Shorrer, Ran I.
Gonczarowski, Yannai A.
author_facet Fish, Sara
Shephard, Julia
Li, Minkai
Shorrer, Ran I.
Gonczarowski, Yannai A.
contents We develop evaluation methods for measuring the economic decision-making capabilities and tendencies of LLMs. First, we develop benchmarks derived from key problems in economics -- procurement, scheduling, and pricing -- that test an LLM's ability to learn from the environment in context. Second, we develop the framework of litmus tests, evaluations that quantify an LLM's choice behavior on a stylized decision-making task with multiple conflicting objectives. Each litmus test outputs a litmus score, which quantifies an LLM's tradeoff response, a reliability score, which measures the coherence of an LLM's choice behavior, and a competency score, which measures an LLM's capability at the same task when the conflicting objectives are replaced by a single, well-specified objective. Evaluating a broad array of frontier LLMs, we (1) investigate changes in LLM capabilities and tendencies over time, (2) derive economically meaningful insights from the LLMs' choice behavior and chain-of-thought, (3) validate our litmus test framework by testing self-consistency, robustness, and generalizability. Overall, this work provides a foundation for evaluating LLM agents as they are further integrated into economic decision-making.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18825
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
Fish, Sara
Shephard, Julia
Li, Minkai
Shorrer, Ran I.
Gonczarowski, Yannai A.
Artificial Intelligence
Computation and Language
Computer Science and Game Theory
We develop evaluation methods for measuring the economic decision-making capabilities and tendencies of LLMs. First, we develop benchmarks derived from key problems in economics -- procurement, scheduling, and pricing -- that test an LLM's ability to learn from the environment in context. Second, we develop the framework of litmus tests, evaluations that quantify an LLM's choice behavior on a stylized decision-making task with multiple conflicting objectives. Each litmus test outputs a litmus score, which quantifies an LLM's tradeoff response, a reliability score, which measures the coherence of an LLM's choice behavior, and a competency score, which measures an LLM's capability at the same task when the conflicting objectives are replaced by a single, well-specified objective. Evaluating a broad array of frontier LLMs, we (1) investigate changes in LLM capabilities and tendencies over time, (2) derive economically meaningful insights from the LLMs' choice behavior and chain-of-thought, (3) validate our litmus test framework by testing self-consistency, robustness, and generalizability. Overall, this work provides a foundation for evaluating LLM agents as they are further integrated into economic decision-making.
title EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
topic Artificial Intelligence
Computation and Language
Computer Science and Game Theory
url https://arxiv.org/abs/2503.18825