FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Fengbin, Ng, Xiang Yao, Liu, Ziyang, Liu, Chang, Zeng, Xianwei, Wang, Chao, Tan, Tianhui, Yao, Xuan, Shao, Pengyang, Xu, Min, Wang, Zixuan, Wang, Jing, Lin, Xin, Li, Junfeng, Zhu, Jingxian, Zhang, Yang, Wang, Wenjie, Feng, Fuli, Hong, Richang, Luan, Huanbo, Huang, Ke-Wei, Chua, Tat-Seng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914237143580672
author Zhu, Fengbin
Ng, Xiang Yao
Liu, Ziyang
Liu, Chang
Zeng, Xianwei
Wang, Chao
Tan, Tianhui
Yao, Xuan
Shao, Pengyang
Xu, Min
Wang, Zixuan
Wang, Jing
Lin, Xin
Li, Junfeng
Zhu, Jingxian
Zhang, Yang
Wang, Wenjie
Feng, Fuli
Hong, Richang
Luan, Huanbo
Huang, Ke-Wei
Chua, Tat-Seng
author_facet Zhu, Fengbin
Ng, Xiang Yao
Liu, Ziyang
Liu, Chang
Zeng, Xianwei
Wang, Chao
Tan, Tianhui
Yao, Xuan
Shao, Pengyang
Xu, Min
Wang, Zixuan
Wang, Jing
Lin, Xin
Li, Junfeng
Zhu, Jingxian
Zhang, Yang
Wang, Wenjie
Feng, Fuli
Hong, Richang
Luan, Huanbo
Huang, Ke-Wei
Chua, Tat-Seng
contents Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks. However, existing literature lacks a rigorous and systematic evaluation of DR Agent's capabilities in critical research analysis. To address this gap, we first propose HisRubric, a novel evaluation framework with a hierarchical analytical structure and a fine-grained grading rubric for rigorously assessing DR agents' capabilities in corporate financial analysis. This framework mirrors the professional analyst's workflow, progressing from data recognition to metric calculation, and finally to strategic summarization and interpretation. Built on this framework, we construct a FinDeepResearch benchmark that comprises 64 listed companies from 8 financial markets across 4 languages, encompassing a total of 15,808 grading items. We further conduct extensive experiments on the FinDeepResearch using 16 representative methods, including 6 DR agents, 5 LLMs equipped with both deep reasoning and search capabilities, and 5 LLMs with deep reasoning capabilities only. The results reveal the strengths and limitations of these approaches across diverse capabilities, financial markets, and languages, offering valuable insights for future research and development. The benchmark and evaluation code is publicly available at https://OpenFinArena.com/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
Zhu, Fengbin
Ng, Xiang Yao
Liu, Ziyang
Liu, Chang
Zeng, Xianwei
Wang, Chao
Tan, Tianhui
Yao, Xuan
Shao, Pengyang
Xu, Min
Wang, Zixuan
Wang, Jing
Lin, Xin
Li, Junfeng
Zhu, Jingxian
Zhang, Yang
Wang, Wenjie
Feng, Fuli
Hong, Richang
Luan, Huanbo
Huang, Ke-Wei
Chua, Tat-Seng
Computation and Language
Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks. However, existing literature lacks a rigorous and systematic evaluation of DR Agent's capabilities in critical research analysis. To address this gap, we first propose HisRubric, a novel evaluation framework with a hierarchical analytical structure and a fine-grained grading rubric for rigorously assessing DR agents' capabilities in corporate financial analysis. This framework mirrors the professional analyst's workflow, progressing from data recognition to metric calculation, and finally to strategic summarization and interpretation. Built on this framework, we construct a FinDeepResearch benchmark that comprises 64 listed companies from 8 financial markets across 4 languages, encompassing a total of 15,808 grading items. We further conduct extensive experiments on the FinDeepResearch using 16 representative methods, including 6 DR agents, 5 LLMs equipped with both deep reasoning and search capabilities, and 5 LLMs with deep reasoning capabilities only. The results reveal the strengths and limitations of these approaches across diverse capabilities, financial markets, and languages, offering valuable insights for future research and development. The benchmark and evaluation code is publicly available at https://OpenFinArena.com/.
title FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
topic Computation and Language
url https://arxiv.org/abs/2510.13936