When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917049922486272 |
|---|---|
| author | Qian, Lingfei Peng, Xueqing Wang, Yan Zhang, Vincent Jim He, Huan Smith, Hanley Han, Yi He, Yueru Li, Haohang Cao, Yupeng Yu, Yangyang Lopez-Lira, Alejandro Lu, Peng Nie, Jian-Yun Xiong, Guojun Huang, Jimin Ananiadou, Sophia |
| author_facet | Qian, Lingfei Peng, Xueqing Wang, Yan Zhang, Vincent Jim He, Huan Smith, Hanley Han, Yi He, Yueru Li, Haohang Cao, Yupeng Yu, Yangyang Lopez-Lira, Alejandro Lu, Peng Nie, Jian-Yun Xiong, Guojun Huang, Jimin Ananiadou, Sophia |
| contents | Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_11695 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents Qian, Lingfei Peng, Xueqing Wang, Yan Zhang, Vincent Jim He, Huan Smith, Hanley Han, Yi He, Yueru Li, Haohang Cao, Yupeng Yu, Yangyang Lopez-Lira, Alejandro Lu, Peng Nie, Jian-Yun Xiong, Guojun Huang, Jimin Ananiadou, Sophia Computation and Language Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies test models instead of agents, cover limited periods and assets, and rely on unverified data. To address these gaps, we introduce Agent Market Arena (AMA), the first lifelong, real-time benchmark for evaluating LLM-based trading agents across multiple markets. AMA integrates verified trading data, expert-checked news, and diverse agent architectures within a unified trading framework, enabling fair and continuous comparison under real conditions. It implements four agents, including InvestorAgent as a single-agent baseline, TradeAgent and HedgeFundAgent with different risk styles, and DeepFundAgent with memory-based reasoning, and evaluates them across GPT-4o, GPT-4.1, Claude-3.5-haiku, Claude-sonnet-4, and Gemini-2.0-flash. Live experiments on both cryptocurrency and stock markets demonstrate that agent frameworks display markedly distinct behavioral patterns, spanning from aggressive risk-taking to conservative decision-making, whereas model backbones contribute less to outcome variation. AMA thus establishes a foundation for rigorous, reproducible, and continuously evolving evaluation of financial reasoning and trading intelligence in LLM-based agents. |
| title | When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.11695 |