When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Lingfei, Peng, Xueqing, Wang, Yan, Zhang, Vincent Jim, He, Huan, Smith, Hanley, Han, Yi, He, Yueru, Li, Haohang, Cao, Yupeng, Yu, Yangyang, Lopez-Lira, Alejandro, Lu, Peng, Nie, Jian-Yun, Xiong, Guojun, Huang, Jimin, Ananiadou, Sophia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917049922486272
author Qian, Lingfei
Peng, Xueqing
Wang, Yan
Zhang, Vincent Jim
He, Huan
Smith, Hanley
Han, Yi
He, Yueru
Li, Haohang
Cao, Yupeng
Yu, Yangyang
Lopez-Lira, Alejandro
Lu, Peng
Nie, Jian-Yun
Xiong, Guojun
Huang, Jimin
Ananiadou, Sophia
author_facet Qian, Lingfei
Peng, Xueqing
Wang, Yan
Zhang, Vincent Jim
He, Huan
Smith, Hanley
Han, Yi
He, Yueru
Li, Haohang
Cao, Yupeng
Yu, Yangyang
Lopez-Lira, Alejandro
Lu, Peng
Nie, Jian-Yun
Xiong, Guojun
Huang, Jimin
Ananiadou, Sophia
contents Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11695
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
Qian, Lingfei
Peng, Xueqing
Wang, Yan
Zhang, Vincent Jim
He, Huan
Smith, Hanley
Han, Yi
He, Yueru
Li, Haohang
Cao, Yupeng
Yu, Yangyang
Lopez-Lira, Alejandro
Lu, Peng
Nie, Jian-Yun
Xiong, Guojun
Huang, Jimin
Ananiadou, Sophia
Computation and Language
Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents.
title When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
topic Computation and Language
url https://arxiv.org/abs/2510.11695