FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913074142773248 |
|---|---|
| author | Xie, Zhuohan Orel, Daniil Thareja, Rushil Sahnan, Dhruv Madmoun, Hachem Zhang, Fan Banerjee, Debopriyo Georgiev, Georgi Peng, Xueqing Qian, Lingfei Huang, Jimin Su, Jinyan Singh, Aaryamonvikram Xing, Rui Elbadry, Rania Xu, Chen Li, Haonan Koto, Fajri Koychev, Ivan Chakraborty, Tanmoy Wang, Yuxia Lahlou, Salem Stoyanov, Veselin Ananiadou, Sophia Nakov, Preslav |
| author_facet | Xie, Zhuohan Orel, Daniil Thareja, Rushil Sahnan, Dhruv Madmoun, Hachem Zhang, Fan Banerjee, Debopriyo Georgiev, Georgi Peng, Xueqing Qian, Lingfei Huang, Jimin Su, Jinyan Singh, Aaryamonvikram Xing, Rui Elbadry, Rania Xu, Chen Li, Haonan Koto, Fajri Koychev, Ivan Chakraborty, Tanmoy Wang, Yuxia Lahlou, Salem Stoyanov, Veselin Ananiadou, Sophia Nakov, Preslav |
| contents | Multi-step symbolic reasoning is essential for robust financial analysis; yet, current benchmarks largely overlook this capability. Existing datasets such as FinQA and ConvFinQA emphasize final numerical answers while neglecting the intermediate reasoning steps required for transparency and verification. To address this gap, we introduce FinChain, the first benchmark specifically designed for verifiable Chain-of-Thought evaluation in finance. FinChain spans 58 topics across 12 financial domains, each represented by parameterized symbolic templates with executable Python code that enable fully machine-verifiable reasoning and scalable, contamination-free data generation. To assess reasoning capacity, we propose CHAINEVAL, a dynamic alignment measure that jointly evaluates both the final-answer correctness and the step-level reasoning consistency. Our evaluation of 26 leading LLMs reveals that even frontier LLMs exhibit clear limitations in symbolic financial reasoning, while domain-adapted and math-enhanced fine-tuned models can substantially narrow this gap. Overall, FinChain exposes persistent weaknesses in multi-step financial reasoning and provides a foundation for developing trustworthy, interpretable, and verifiable financial AI. This project is available at https://github.com/mbzuai-nlp/finchain.git. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_02515 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning Xie, Zhuohan Orel, Daniil Thareja, Rushil Sahnan, Dhruv Madmoun, Hachem Zhang, Fan Banerjee, Debopriyo Georgiev, Georgi Peng, Xueqing Qian, Lingfei Huang, Jimin Su, Jinyan Singh, Aaryamonvikram Xing, Rui Elbadry, Rania Xu, Chen Li, Haonan Koto, Fajri Koychev, Ivan Chakraborty, Tanmoy Wang, Yuxia Lahlou, Salem Stoyanov, Veselin Ananiadou, Sophia Nakov, Preslav Computation and Language Artificial Intelligence Machine Learning Multi-step symbolic reasoning is essential for robust financial analysis; yet, current benchmarks largely overlook this capability. Existing datasets such as FinQA and ConvFinQA emphasize final numerical answers while neglecting the intermediate reasoning steps required for transparency and verification. To address this gap, we introduce FinChain, the first benchmark specifically designed for verifiable Chain-of-Thought evaluation in finance. FinChain spans 58 topics across 12 financial domains, each represented by parameterized symbolic templates with executable Python code that enable fully machine-verifiable reasoning and scalable, contamination-free data generation. To assess reasoning capacity, we propose CHAINEVAL, a dynamic alignment measure that jointly evaluates both the final-answer correctness and the step-level reasoning consistency. Our evaluation of 26 leading LLMs reveals that even frontier LLMs exhibit clear limitations in symbolic financial reasoning, while domain-adapted and math-enhanced fine-tuned models can substantially narrow this gap. Overall, FinChain exposes persistent weaknesses in multi-step financial reasoning and provides a foundation for developing trustworthy, interpretable, and verifiable financial AI. This project is available at https://github.com/mbzuai-nlp/finchain.git. |
| title | FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2506.02515 |