FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Zhuohan, Orel, Daniil, Thareja, Rushil, Sahnan, Dhruv, Madmoun, Hachem, Zhang, Fan, Banerjee, Debopriyo, Georgiev, Georgi, Peng, Xueqing, Qian, Lingfei, Huang, Jimin, Su, Jinyan, Singh, Aaryamonvikram, Xing, Rui, Elbadry, Rania, Xu, Chen, Li, Haonan, Koto, Fajri, Koychev, Ivan, Chakraborty, Tanmoy, Wang, Yuxia, Lahlou, Salem, Stoyanov, Veselin, Ananiadou, Sophia, Nakov, Preslav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913074142773248
author Xie, Zhuohan
Orel, Daniil
Thareja, Rushil
Sahnan, Dhruv
Madmoun, Hachem
Zhang, Fan
Banerjee, Debopriyo
Georgiev, Georgi
Peng, Xueqing
Qian, Lingfei
Huang, Jimin
Su, Jinyan
Singh, Aaryamonvikram
Xing, Rui
Elbadry, Rania
Xu, Chen
Li, Haonan
Koto, Fajri
Koychev, Ivan
Chakraborty, Tanmoy
Wang, Yuxia
Lahlou, Salem
Stoyanov, Veselin
Ananiadou, Sophia
Nakov, Preslav
author_facet Xie, Zhuohan
Orel, Daniil
Thareja, Rushil
Sahnan, Dhruv
Madmoun, Hachem
Zhang, Fan
Banerjee, Debopriyo
Georgiev, Georgi
Peng, Xueqing
Qian, Lingfei
Huang, Jimin
Su, Jinyan
Singh, Aaryamonvikram
Xing, Rui
Elbadry, Rania
Xu, Chen
Li, Haonan
Koto, Fajri
Koychev, Ivan
Chakraborty, Tanmoy
Wang, Yuxia
Lahlou, Salem
Stoyanov, Veselin
Ananiadou, Sophia
Nakov, Preslav
contents Multi-step symbolic reasoning is essential for robust financial analysis; yet, current benchmarks largely overlook this capability. Existing datasets such as FinQA and ConvFinQA emphasize final numerical answers while neglecting the intermediate reasoning steps required for transparency and verification. To address this gap, we introduce FinChain, the first benchmark specifically designed for verifiable Chain-of-Thought evaluation in finance. FinChain spans 58 topics across 12 financial domains, each represented by parameterized symbolic templates with executable Python code that enable fully machine-verifiable reasoning and scalable, contamination-free data generation. To assess reasoning capacity, we propose CHAINEVAL, a dynamic alignment measure that jointly evaluates both the final-answer correctness and the step-level reasoning consistency. Our evaluation of 26 leading LLMs reveals that even frontier LLMs exhibit clear limitations in symbolic financial reasoning, while domain-adapted and math-enhanced fine-tuned models can substantially narrow this gap. Overall, FinChain exposes persistent weaknesses in multi-step financial reasoning and provides a foundation for developing trustworthy, interpretable, and verifiable financial AI. This project is available at https://github.com/mbzuai-nlp/finchain.git.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02515
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
Xie, Zhuohan
Orel, Daniil
Thareja, Rushil
Sahnan, Dhruv
Madmoun, Hachem
Zhang, Fan
Banerjee, Debopriyo
Georgiev, Georgi
Peng, Xueqing
Qian, Lingfei
Huang, Jimin
Su, Jinyan
Singh, Aaryamonvikram
Xing, Rui
Elbadry, Rania
Xu, Chen
Li, Haonan
Koto, Fajri
Koychev, Ivan
Chakraborty, Tanmoy
Wang, Yuxia
Lahlou, Salem
Stoyanov, Veselin
Ananiadou, Sophia
Nakov, Preslav
Computation and Language
Artificial Intelligence
Machine Learning
Multi-step symbolic reasoning is essential for robust financial analysis; yet, current benchmarks largely overlook this capability. Existing datasets such as FinQA and ConvFinQA emphasize final numerical answers while neglecting the intermediate reasoning steps required for transparency and verification. To address this gap, we introduce FinChain, the first benchmark specifically designed for verifiable Chain-of-Thought evaluation in finance. FinChain spans 58 topics across 12 financial domains, each represented by parameterized symbolic templates with executable Python code that enable fully machine-verifiable reasoning and scalable, contamination-free data generation. To assess reasoning capacity, we propose CHAINEVAL, a dynamic alignment measure that jointly evaluates both the final-answer correctness and the step-level reasoning consistency. Our evaluation of 26 leading LLMs reveals that even frontier LLMs exhibit clear limitations in symbolic financial reasoning, while domain-adapted and math-enhanced fine-tuned models can substantially narrow this gap. Overall, FinChain exposes persistent weaknesses in multi-step financial reasoning and provides a foundation for developing trustworthy, interpretable, and verifiable financial AI. This project is available at https://github.com/mbzuai-nlp/finchain.git.
title FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.02515