Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Yanxu, Yao, Zijun, Liu, Yantao, Xin, Amy, Ye, Jin, Yu, Jianing, Hou, Lei, Li, Juanzi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2510.02209
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910037427879936
author Chen, Yanxu
Yao, Zijun
Liu, Yantao
Xin, Amy
Ye, Jin
Yu, Jianing
Hou, Lei
Li, Juanzi
author_facet Chen, Yanxu
Yao, Zijun
Liu, Yantao
Xin, Amy
Ye, Jin
Yu, Jianing
Hou, Lei
Li, Juanzi
contents Large language models (LLMs) demonstrate strong potential as autonomous agents, with promising capabilities in reasoning, tool use, and sequential decision-making. While prior benchmarks have evaluated LLM agents in various domains, the financial domain remains underexplored, despite its significant economic value and complex reasoning requirements. Most existing financial benchmarks focus on static question-answering, failing to capture the dynamics of real-market trading. To address this gap, we introduce STOCKBENCH, a contamination-free benchmark designed to evaluate LLM agents in realistic, multi-month stock trading environments. Agents receive daily market signals -- including prices, fundamentals, and news -- and make sequential buy, sell, or hold decisions. Performance is measured using financial metrics such as cumulative return, maximum drawdown, and the Sortino ratio, capturing both profitability and risk management. We evaluate a wide range of state-of-the-art proprietary and open-source LLMs. Surprisingly, most models struggle to outperform the simple buy-and-hold baseline, while some models demonstrate the potential to achieve higher returns and stronger risk management. These findings highlight both the challenges and opportunities of LLM-based trading agents, showing that strong performance on static financial question-answering do not necessarily translate into effective trading behavior. We release STOCKBENCH as an open-source benchmark to enable future research on LLM-driven financial agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02209
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
Chen, Yanxu
Yao, Zijun
Liu, Yantao
Xin, Amy
Ye, Jin
Yu, Jianing
Hou, Lei
Li, Juanzi
Machine Learning
Computation and Language
Large language models (LLMs) demonstrate strong potential as autonomous agents, with promising capabilities in reasoning, tool use, and sequential decision-making. While prior benchmarks have evaluated LLM agents in various domains, the financial domain remains underexplored, despite its significant economic value and complex reasoning requirements. Most existing financial benchmarks focus on static question-answering, failing to capture the dynamics of real-market trading. To address this gap, we introduce STOCKBENCH, a contamination-free benchmark designed to evaluate LLM agents in realistic, multi-month stock trading environments. Agents receive daily market signals -- including prices, fundamentals, and news -- and make sequential buy, sell, or hold decisions. Performance is measured using financial metrics such as cumulative return, maximum drawdown, and the Sortino ratio, capturing both profitability and risk management. We evaluate a wide range of state-of-the-art proprietary and open-source LLMs. Surprisingly, most models struggle to outperform the simple buy-and-hold baseline, while some models demonstrate the potential to achieve higher returns and stronger risk management. These findings highlight both the challenges and opportunities of LLM-based trading agents, showing that strong performance on static financial question-answering do not necessarily translate into effective trading behavior. We release STOCKBENCH as an open-source benchmark to enable future research on LLM-driven financial agents.
title StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2510.02209