MarketBench: Evaluating AI Agents as Market Participants

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fradkin, Andrey, Krishnan, Rohit
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911625202630656
author Fradkin, Andrey
Krishnan, Rohit
author_facet Fradkin, Andrey
Krishnan, Rohit
contents Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to successfully complete a task and the cost of doing so. We propose MarketBench, a benchmark for assessing whether AI agents have these capabilities. We use a 93-task subset of SWE-bench Lite, a software engineering benchmark, with six recently released LLMs as a demonstration. These LLMs are miscalibrated on both success probability and token usage, and auctions built from these self-reports diverge from a full-information allocation. A follow-up intervention where we add information about capabilities from prior experiments to the context improves calibration, but only modestly narrows the gap to a full-information benchmark. We also document the performance of a market-based scaffolding with these LLMs. Our results point to self-assessment as a key bottleneck for market-style coordination of AI agents.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23897
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MarketBench: Evaluating AI Agents as Market Participants
Fradkin, Andrey
Krishnan, Rohit
Artificial Intelligence
General Economics
Economics
Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to successfully complete a task and the cost of doing so. We propose MarketBench, a benchmark for assessing whether AI agents have these capabilities. We use a 93-task subset of SWE-bench Lite, a software engineering benchmark, with six recently released LLMs as a demonstration. These LLMs are miscalibrated on both success probability and token usage, and auctions built from these self-reports diverge from a full-information allocation. A follow-up intervention where we add information about capabilities from prior experiments to the context improves calibration, but only modestly narrows the gap to a full-information benchmark. We also document the performance of a market-based scaffolding with these LLMs. Our results point to self-assessment as a key bottleneck for market-style coordination of AI agents.
title MarketBench: Evaluating AI Agents as Market Participants
topic Artificial Intelligence
General Economics
Economics
url https://arxiv.org/abs/2604.23897