Understanding Financial Reasoning in AI: A Multimodal Benchmark and Error Learning Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Shuangyan, Peng, Haizhou, Xu, Jiachen, Liu, Chunhou, Giurcuaneanu, Ciprian Doru, Liu, Jiamou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
by: Deng, Shuangyan, et al.
Published: (2025)
by: Deng, Shuangyan, et al.
Published: (2025)
Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
From Narrative to Action: A Hierarchical LLM-Agent Framework for Human Mobility Generation
by: Li, Qiumeng, et al.
Published: (2025)
by: Li, Qiumeng, et al.
Published: (2025)
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
by: Hu, He, et al.
Published: (2026)
by: Hu, He, et al.
Published: (2026)
Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
by: Liu, Jiachen, et al.
Published: (2026)
by: Liu, Jiachen, et al.
Published: (2026)
Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning
by: Yin, Hang, et al.
Published: (2024)
by: Yin, Hang, et al.
Published: (2024)
ChatLogic: Integrating Logic Programming with Large Language Models for Multi-Step Reasoning
by: Wang, Zhongsheng, et al.
Published: (2024)
by: Wang, Zhongsheng, et al.
Published: (2024)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
When Does Multimodal Learning Help in Healthcare? A Benchmark on EHR and Chest X-Ray Fusion
by: Yin, Kejing, et al.
Published: (2026)
by: Yin, Kejing, et al.
Published: (2026)
Benchmarking Zero-Shot Reasoning Approaches for Error Detection in Solidity Smart Contracts
by: Sardenberg, Eduardo, et al.
Published: (2026)
by: Sardenberg, Eduardo, et al.
Published: (2026)
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
by: Yakun, Cui, et al.
Published: (2025)
by: Yakun, Cui, et al.
Published: (2025)
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
by: Zhao, Boxiang, et al.
Published: (2026)
by: Zhao, Boxiang, et al.
Published: (2026)
Revisiting, Benchmarking and Understanding Unsupervised Graph Domain Adaptation
by: Liu, Meihan, et al.
Published: (2024)
by: Liu, Meihan, et al.
Published: (2024)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs
by: Ning, Yansong, et al.
Published: (2026)
by: Ning, Yansong, et al.
Published: (2026)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Stepwise Reasoning Error Disruption Attack of LLMs
by: Peng, Jingyu, et al.
Published: (2024)
by: Peng, Jingyu, et al.
Published: (2024)
Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study
by: Yao, Xuan, et al.
Published: (2025)
by: Yao, Xuan, et al.
Published: (2025)
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
by: Bao, Qiming, et al.
Published: (2022)
by: Bao, Qiming, et al.
Published: (2022)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety
by: Liu, Junliang, et al.
Published: (2025)
by: Liu, Junliang, et al.
Published: (2025)
Dimensions of Vulnerability in Visual Working Memory: An AI-Driven Approach to Perceptual Comparison
by: Cao, Yuang, et al.
Published: (2025)
by: Cao, Yuang, et al.
Published: (2025)
S2ED: From Story to Executable Descriptions for Consistency-Aware Story Illustration
by: Yin, Sijing, et al.
Published: (2026)
by: Yin, Sijing, et al.
Published: (2026)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
Situational-Constrained Sequential Resources Allocation via Reinforcement Learning
by: Zhang, Libo, et al.
Published: (2025)
by: Zhang, Libo, et al.
Published: (2025)
Inferring Reward Machines and Transition Machines from Partially Observable Markov Decision Processes
by: Wu, Yuly, et al.
Published: (2025)
by: Wu, Yuly, et al.
Published: (2025)
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
by: Zhou, Yitong, et al.
Published: (2025)
by: Zhou, Yitong, et al.
Published: (2025)
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
by: Zhao, Zehua, et al.
Published: (2025)
by: Zhao, Zehua, et al.
Published: (2025)
Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
by: Liu, Weichen, et al.
Published: (2025)
by: Liu, Weichen, et al.
Published: (2025)
Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
by: Xu, Kaishuai, et al.
Published: (2024)
by: Xu, Kaishuai, et al.
Published: (2024)
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning
by: Xia, Haizhou
Published: (2026)
by: Xia, Haizhou
Published: (2026)
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning
by: Bao, Qiming, et al.
Published: (2023)
by: Bao, Qiming, et al.
Published: (2023)
CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution
by: Xu, Ruiyang, et al.
Published: (2024)
by: Xu, Ruiyang, et al.
Published: (2024)
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
by: Yue, Xiang, et al.
Published: (2023)
by: Yue, Xiang, et al.
Published: (2023)
Radiology's Last Exam (RadLE): Benchmarking Frontier Multimodal AI Against Human Experts and a Taxonomy of Visual Reasoning Errors in Radiology
by: Datta, Suvrankar, et al.
Published: (2025)
by: Datta, Suvrankar, et al.
Published: (2025)
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
by: Wei, Jingxuan, et al.
Published: (2025)
by: Wei, Jingxuan, et al.
Published: (2025)
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
by: Zhang, Xiyuan, et al.
Published: (2026)
by: Zhang, Xiyuan, et al.
Published: (2026)
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
by: Yan, Qianqi, et al.
Published: (2025)
by: Yan, Qianqi, et al.
Published: (2025)
Similar Items
-
FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
by: Deng, Shuangyan, et al.
Published: (2025) -
Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference
by: Li, Haoyuan, et al.
Published: (2025) -
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
by: Liu, Rui, et al.
Published: (2024) -
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
by: Xu, Zhaopan, et al.
Published: (2025) -
From Narrative to Action: A Hierarchical LLM-Agent Framework for Human Mobility Generation
by: Li, Qiumeng, et al.
Published: (2025)