FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Zichen, E, Haihong, Ma, Ziyan, He, Haoyang, Liu, Jiacheng, Yang, Zhongjun, Rong, Zihua, Li, Rongjin, Ji, Kun, Huang, Qing, Hu, Xinyang, Liu, Yang, Zheng, Qianhe
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913977520357376
author Tang, Zichen
E, Haihong
Ma, Ziyan
He, Haoyang
Liu, Jiacheng
Yang, Zhongjun
Rong, Zihua
Li, Rongjin
Ji, Kun
Huang, Qing
Hu, Xinyang
Liu, Yang
Zheng, Qianhe
author_facet Tang, Zichen
E, Haihong
Ma, Ziyan
He, Haoyang
Liu, Jiacheng
Yang, Zhongjun
Rong, Zihua
Li, Rongjin
Ji, Kun
Huang, Qing
Hu, Xinyang
Liu, Yang
Zheng, Qianhe
contents We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compared to existing benchmarks, our work provides three key advancements. (1) Credibility: We update 15.6% of the questions from four public datasets, annotating 908 new questions with detailed Python solutions and rigorously refining evaluation standards. This enables an accurate assessment of the reasoning improvements of LRMs. (2) Comprehensiveness: FinanceReasoning covers 67.8% of financial concepts and formulas, significantly surpassing existing datasets. Additionally, we construct 3,133 Python-formatted functions, which enhances LRMs' financial reasoning capabilities through refined knowledge (e.g., 83.2% $\rightarrow$ 91.6% for GPT-4o). (3) Challenge: Models are required to apply multiple financial formulas for precise numerical reasoning on 238 Hard problems. The best-performing model (i.e., OpenAI o1 with PoT) achieves 89.1% accuracy, yet LRMs still face challenges in numerical precision. We demonstrate that combining Reasoner and Programmer models can effectively enhance LRMs' performance (e.g., 83.2% $\rightarrow$ 87.8% for DeepSeek-R1). Our work paves the way for future research on evaluating and improving LRMs in domain-specific complex reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
Tang, Zichen
E, Haihong
Ma, Ziyan
He, Haoyang
Liu, Jiacheng
Yang, Zhongjun
Rong, Zihua
Li, Rongjin
Ji, Kun
Huang, Qing
Hu, Xinyang
Liu, Yang
Zheng, Qianhe
Computation and Language
Computational Engineering, Finance, and Science
We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compared to existing benchmarks, our work provides three key advancements. (1) Credibility: We update 15.6% of the questions from four public datasets, annotating 908 new questions with detailed Python solutions and rigorously refining evaluation standards. This enables an accurate assessment of the reasoning improvements of LRMs. (2) Comprehensiveness: FinanceReasoning covers 67.8% of financial concepts and formulas, significantly surpassing existing datasets. Additionally, we construct 3,133 Python-formatted functions, which enhances LRMs' financial reasoning capabilities through refined knowledge (e.g., 83.2% $\rightarrow$ 91.6% for GPT-4o). (3) Challenge: Models are required to apply multiple financial formulas for precise numerical reasoning on 238 Hard problems. The best-performing model (i.e., OpenAI o1 with PoT) achieves 89.1% accuracy, yet LRMs still face challenges in numerical precision. We demonstrate that combining Reasoner and Programmer models can effectively enhance LRMs' performance (e.g., 83.2% $\rightarrow$ 87.8% for DeepSeek-R1). Our work paves the way for future research on evaluating and improving LRMs in domain-specific complex reasoning tasks.
title FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
topic Computation and Language
Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2506.05828