Saved in:
Bibliographic Details
Main Authors: Li, Changlun, Shi, Yao, Wang, Chen, Duan, Qiqi, Ruan, Runke, Huang, Weijie, Long, Haonan, Huang, Lijun, Tang, Nan, Luo, Yuyu
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.11065
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908590431797248
author Li, Changlun
Shi, Yao
Wang, Chen
Duan, Qiqi
Ruan, Runke
Huang, Weijie
Long, Haonan
Huang, Lijun
Tang, Nan
Luo, Yuyu
author_facet Li, Changlun
Shi, Yao
Wang, Chen
Duan, Qiqi
Ruan, Runke
Huang, Weijie
Long, Haonan
Huang, Lijun
Tang, Nan
Luo, Yuyu
contents Large Language Models (LLMs) have demonstrated notable capabilities across financial tasks, including financial report summarization, earnings call transcript analysis, and asset classification. However, their real-world effectiveness in managing complex fund investment remains inadequately assessed. A fundamental limitation of existing benchmarks for evaluating LLM-driven trading strategies is their reliance on historical back-testing, inadvertently enabling LLMs to "time travel"-leveraging future information embedded in their training corpora, thus resulting in possible information leakage and overly optimistic performance estimates. To address this issue, we introduce DeepFund, a live fund benchmark tool designed to rigorously evaluate LLM in real-time market conditions. Utilizing a multi-agent architecture, DeepFund connects directly with real-time stock market data-specifically data published after each model pretraining cutoff-to ensure fair and leakage-free evaluations. Empirical tests on nine flagship LLMs from leading global institutions across multiple investment dimensions-including ticker-level analysis, investment decision-making, portfolio management, and risk control-reveal significant practical challenges. Notably, even cutting-edge models such as DeepSeek-V3 and Claude-3.7-Sonnet incur net trading losses within DeepFund real-time evaluation environment, underscoring the present limitations of LLMs for active fund management. Our code is available at https://github.com/HKUSTDial/DeepFund.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11065
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking
Li, Changlun
Shi, Yao
Wang, Chen
Duan, Qiqi
Ruan, Runke
Huang, Weijie
Long, Haonan
Huang, Lijun
Tang, Nan
Luo, Yuyu
Computational Engineering, Finance, and Science
Artificial Intelligence
Multiagent Systems
Large Language Models (LLMs) have demonstrated notable capabilities across financial tasks, including financial report summarization, earnings call transcript analysis, and asset classification. However, their real-world effectiveness in managing complex fund investment remains inadequately assessed. A fundamental limitation of existing benchmarks for evaluating LLM-driven trading strategies is their reliance on historical back-testing, inadvertently enabling LLMs to "time travel"-leveraging future information embedded in their training corpora, thus resulting in possible information leakage and overly optimistic performance estimates. To address this issue, we introduce DeepFund, a live fund benchmark tool designed to rigorously evaluate LLM in real-time market conditions. Utilizing a multi-agent architecture, DeepFund connects directly with real-time stock market data-specifically data published after each model pretraining cutoff-to ensure fair and leakage-free evaluations. Empirical tests on nine flagship LLMs from leading global institutions across multiple investment dimensions-including ticker-level analysis, investment decision-making, portfolio management, and risk control-reveal significant practical challenges. Notably, even cutting-edge models such as DeepSeek-V3 and Claude-3.7-Sonnet incur net trading losses within DeepFund real-time evaluation environment, underscoring the present limitations of LLMs for active fund management. Our code is available at https://github.com/HKUSTDial/DeepFund.
title Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking
topic Computational Engineering, Finance, and Science
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2505.11065