FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Junyu, Kou, Zhizhuo, Yang, Liming, Luo, Xiao, Huang, Jinsheng, Xiao, Zhiping, Peng, Jingshu, Liu, Chengzhong, Ji, Jiaming, Liu, Xuanzhe, Han, Sirui, Zhang, Ming, Guo, Yike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automate Strategy Finding with LLM in Quant Investment
von: Kou, Zhizhuo, et al.
Veröffentlicht: (2024)
von: Kou, Zhizhuo, et al.
Veröffentlicht: (2024)
Attention Bootstrapping for Multi-Modal Test-Time Adaptation
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
Sparse Causal Discovery with Generative Intervention for Unsupervised Graph Domain Adaptation
von: Luo, Junyu, et al.
Veröffentlicht: (2025)
von: Luo, Junyu, et al.
Veröffentlicht: (2025)
MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark
von: Yakun, Cui, et al.
Veröffentlicht: (2026)
von: Yakun, Cui, et al.
Veröffentlicht: (2026)
BAMAS: Structuring Budget-Aware Multi-Agent Systems
von: Yang, Liming, et al.
Veröffentlicht: (2025)
von: Yang, Liming, et al.
Veröffentlicht: (2025)
Semi-supervised Fine-tuning for Large Language Models
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
von: Wang, Dannong, et al.
Veröffentlicht: (2025)
von: Wang, Dannong, et al.
Veröffentlicht: (2025)
RW-TTT: Batched Serving for Request-Owned Test-Time Training State
von: Yang, Jian, et al.
Veröffentlicht: (2026)
von: Yang, Jian, et al.
Veröffentlicht: (2026)
MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2025)
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
von: Gan, Ziliang, et al.
Veröffentlicht: (2024)
Rank and Align: Towards Effective Source-free Graph Domain Adaptation
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
Benchmarking Multi-National Value Alignment for Large Language Models
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
von: Shi, Weijie, et al.
Veröffentlicht: (2025)
FinRL Contests: Benchmarking Data-driven Financial Reinforcement Learning Agents
von: Wang, Keyi, et al.
Veröffentlicht: (2025)
von: Wang, Keyi, et al.
Veröffentlicht: (2025)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
FinAnchor: Aligned Multi-Model Representations for Financial Prediction
von: He, Zirui, et al.
Veröffentlicht: (2026)
von: He, Zirui, et al.
Veröffentlicht: (2026)
GALA: Graph Diffusion-based Alignment with Jigsaw for Source-free Domain Adaptation
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
Cross-Domain Diffusion with Progressive Alignment for Efficient Adaptive Retrieval
von: Luo, Junyu, et al.
Veröffentlicht: (2025)
von: Luo, Junyu, et al.
Veröffentlicht: (2025)
FinBen: A Holistic Financial Benchmark for Large Language Models
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
von: Xie, Qianqian, et al.
Veröffentlicht: (2024)
MEME: Modeling the Evolutionary Modes of Financial Markets
von: Guo, Taian, et al.
Veröffentlicht: (2026)
von: Guo, Taian, et al.
Veröffentlicht: (2026)
FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
Reimagining Legal Fact Verification with GenAI: Toward Effective Human-AI Collaboration
von: Han, Sirui, et al.
Veröffentlicht: (2026)
von: Han, Sirui, et al.
Veröffentlicht: (2026)
FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
von: Deng, Shuangyan, et al.
Veröffentlicht: (2025)
von: Deng, Shuangyan, et al.
Veröffentlicht: (2025)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
von: Agrawal, Yogesh, et al.
Veröffentlicht: (2026)
FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis
von: Xu, Wenyan, et al.
Veröffentlicht: (2025)
von: Xu, Wenyan, et al.
Veröffentlicht: (2025)
MARCO: Meta-Reflection with Cross-Referencing for Code Reasoning
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents
von: Zhu, Zhenghao, et al.
Veröffentlicht: (2025)
von: Zhu, Zhenghao, et al.
Veröffentlicht: (2025)
FinGPT: Open-Source Financial Large Language Models
von: Yang, Hongyang, et al.
Veröffentlicht: (2023)
von: Yang, Hongyang, et al.
Veröffentlicht: (2023)
PGODE: Towards High-quality System Dynamics Modeling
von: Luo, Xiao, et al.
Veröffentlicht: (2023)
von: Luo, Xiao, et al.
Veröffentlicht: (2023)
DocFinQA: A Long-Context Financial Reasoning Dataset
von: Reddy, Varshini, et al.
Veröffentlicht: (2024)
von: Reddy, Varshini, et al.
Veröffentlicht: (2024)
FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reporting
von: Zhu, Yiyun, et al.
Veröffentlicht: (2026)
von: Zhu, Yiyun, et al.
Veröffentlicht: (2026)
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
von: Liu, Yuansen, et al.
Veröffentlicht: (2025)
von: Liu, Yuansen, et al.
Veröffentlicht: (2025)
A Survey of Data-Efficient Graph Learning
von: Ju, Wei, et al.
Veröffentlicht: (2024)
von: Ju, Wei, et al.
Veröffentlicht: (2024)
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
von: Lu, Guilong, et al.
Veröffentlicht: (2025)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
von: Xie, Zhuohan, et al.
Veröffentlicht: (2025)
von: Xie, Zhuohan, et al.
Veröffentlicht: (2025)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
von: Zhan, Weidong, et al.
Veröffentlicht: (2025)
von: Zhan, Weidong, et al.
Veröffentlicht: (2025)
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
von: Cao, Yupeng, et al.
Veröffentlicht: (2025)
Benchmarking Automated Clinical Language Simplification: Dataset, Algorithm, and Evaluation
von: Luo, Junyu, et al.
Veröffentlicht: (2020)
von: Luo, Junyu, et al.
Veröffentlicht: (2020)
Ähnliche Einträge
-
Automate Strategy Finding with LLM in Quant Investment
von: Kou, Zhizhuo, et al.
Veröffentlicht: (2024) -
Attention Bootstrapping for Multi-Modal Test-Time Adaptation
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025) -
Sparse Causal Discovery with Generative Intervention for Unsupervised Graph Domain Adaptation
von: Luo, Junyu, et al.
Veröffentlicht: (2025) -
MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark
von: Yakun, Cui, et al.
Veröffentlicht: (2026) -
BAMAS: Structuring Budget-Aware Multi-Agent Systems
von: Yang, Liming, et al.
Veröffentlicht: (2025)