FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Zhuohan, Orel, Daniil, Thareja, Rushil, Sahnan, Dhruv, Madmoun, Hachem, Zhang, Fan, Banerjee, Debopriyo, Georgiev, Georgi, Peng, Xueqing, Qian, Lingfei, Huang, Jimin, Su, Jinyan, Singh, Aaryamonvikram, Xing, Rui, Elbadry, Rania, Xu, Chen, Li, Haonan, Koto, Fajri, Koychev, Ivan, Chakraborty, Tanmoy, Wang, Yuxia, Lahlou, Salem, Stoyanov, Veselin, Ananiadou, Sophia, Nakov, Preslav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
by: Gull, Ayesha, et al.
Published: (2025)
by: Gull, Ayesha, et al.
Published: (2025)
Communication Enables Cooperation in LLM Agents: A Comparison with Curriculum-Based Approaches
by: Madmoun, Hachem, et al.
Published: (2025)
by: Madmoun, Hachem, et al.
Published: (2025)
The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems
by: Xie, Zhuohan, et al.
Published: (2026)
by: Xie, Zhuohan, et al.
Published: (2026)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
by: Thareja, Rushil, et al.
Published: (2025)
by: Thareja, Rushil, et al.
Published: (2025)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)
by: Sadallah, Abdelrahman, et al.
Published: (2026)
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
by: Laiyk, Nurkhan, et al.
Published: (2025)
by: Laiyk, Nurkhan, et al.
Published: (2025)
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations
by: Elbadry, Rania, et al.
Published: (2026)
by: Elbadry, Rania, et al.
Published: (2026)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
by: Almheiri, Saeed, et al.
Published: (2025)
by: Almheiri, Saeed, et al.
Published: (2025)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia
by: Koto, Fajri
Published: (2024)
by: Koto, Fajri
Published: (2024)
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
by: Ivanov, Petar, et al.
Published: (2023)
by: Ivanov, Petar, et al.
Published: (2023)
YaPO: Learnable Sparse Activation Steering Vectors for Domain Adaptation
by: Bounhar, Abdelaziz, et al.
Published: (2026)
by: Bounhar, Abdelaziz, et al.
Published: (2026)
Adapting Fake News Detection to the Era of Large Language Models
by: Su, Jinyan, et al.
Published: (2023)
by: Su, Jinyan, et al.
Published: (2023)
Corpus Poisoning via Approximate Greedy Gradient Descent
by: Su, Jinyan, et al.
Published: (2024)
by: Su, Jinyan, et al.
Published: (2024)
FIRE: Fact-checking with Iterative Retrieval and Verification
by: Xie, Zhuohan, et al.
Published: (2024)
by: Xie, Zhuohan, et al.
Published: (2024)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
by: Sahnan, Dhruv, et al.
Published: (2026)
by: Sahnan, Dhruv, et al.
Published: (2026)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
by: Tomar, Raj Vardhan, et al.
Published: (2025)
by: Tomar, Raj Vardhan, et al.
Published: (2025)
FMI_SU_Yotkova_Kastreva at SemEval-2026 Task 13: Lightweight Detection of LLM-Generated Code via Stylometric Signals
by: Yotkova, Elitsa, et al.
Published: (2026)
by: Yotkova, Elitsa, et al.
Published: (2026)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Fast or Better? Balancing Accuracy and Cost in Retrieval-Augmented Generation with Flexible User Control
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Post-OCR Text Correction for Bulgarian Historical Documents
by: Beshirov, Angel, et al.
Published: (2024)
by: Beshirov, Angel, et al.
Published: (2024)
Can LLMs Automate Fact-Checking Article Writing?
by: Sahnan, Dhruv, et al.
Published: (2025)
by: Sahnan, Dhruv, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
by: Yari, Amir Hossein, et al.
Published: (2025)
by: Yari, Amir Hossein, et al.
Published: (2025)
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
by: Choukrani, Omar, et al.
Published: (2025)
by: Choukrani, Omar, et al.
Published: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2026)
by: Orel, Daniil, et al.
Published: (2026)
Parallel Tokenizers: Rethinking Vocabulary Design for Cross-Lingual Transfer
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2025)
A Study in Markov Chains, Loop-Erased Random Walk and Loop Soups
by: Gu, Zhuohan
Published: (2024)
by: Gu, Zhuohan
Published: (2024)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
by: Goloburda, Maiya, et al.
Published: (2025)
by: Goloburda, Maiya, et al.
Published: (2025)
Offline and Online KL-Regularized RLHF under Differential Privacy
by: Wu, Yulian, et al.
Published: (2025)
by: Wu, Yulian, et al.
Published: (2025)
MAC: Multi-Agent Constitution Learning
by: Thareja, Rushil, et al.
Published: (2026)
by: Thareja, Rushil, et al.
Published: (2026)
FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering
by: Zhou, Yixi, et al.
Published: (2026)
by: Zhou, Yixi, et al.
Published: (2026)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
by: Togmanov, Mukhammed, et al.
Published: (2025)
by: Togmanov, Mukhammed, et al.
Published: (2025)
Towards More Robust Retrieval-Augmented Generation: Evaluating RAG Under Adversarial Poisoning Attacks
by: Su, Jinyan, et al.
Published: (2024)
by: Su, Jinyan, et al.
Published: (2024)
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
by: Das, Rocktim Jyoti, et al.
Published: (2024)
by: Das, Rocktim Jyoti, et al.
Published: (2024)
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
by: Aziz, Rashad, et al.
Published: (2026)
by: Aziz, Rashad, et al.
Published: (2026)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-Checking
by: Gong, Shuzhi, et al.
Published: (2026)
by: Gong, Shuzhi, et al.
Published: (2026)
Similar Items
-
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
by: Gull, Ayesha, et al.
Published: (2025) -
Communication Enables Cooperation in LLM Agents: A Comparison with Curriculum-Based Approaches
by: Madmoun, Hachem, et al.
Published: (2025) -
The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems
by: Xie, Zhuohan, et al.
Published: (2026) -
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
by: Thareja, Rushil, et al.
Published: (2025) -
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)