MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yakun, Cui, Zhang, Yanting, Lei, Zhu, Xie, Jian, Kou, Zhizhuo, Du, Hang, Zhu, Zhenghao, Han, Sirui
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918280853192704
author Yakun, Cui
Zhang, Yanting
Lei, Zhu
Xie, Jian
Kou, Zhizhuo
Du, Hang
Zhu, Zhenghao
Han, Sirui
author_facet Yakun, Cui
Zhang, Yanting
Lei, Zhu
Xie, Jian
Kou, Zhizhuo
Du, Hang
Zhu, Zhenghao
Han, Sirui
contents The advent of multi-modal language models (MLLMs) has spurred research into their application across various table understanding tasks. However, their performance in credit table understanding (CTU) for financial credit review remains largely unexplored due to the following barriers: low data consistency, high annotation costs stemming from domain-specific knowledge and complex calculations, and evaluation paradigm gaps between benchmark and real-world scenarios. To address these challenges, we introduce MMFCTUB (Multi-Modal Financial Credit Table Understanding Benchmark), a practical benchmark, encompassing more than 7,600 high quality CTU samples across 5 table types. MMFCTUB employ a minimally supervised pipeline that adheres to inter-table constraints and maintains data distributions consistency. The benchmark leverages capacity-driven questions and mask-and-recovery strategy to evaluate models' cross-table structure perception, domain knowledge utilization, and numerical calculation capabilities. Utilizing MMFCTUB, we conduct comprehensive evaluations of both proprietary and open-source MLLMs, revealing their strengths and limitations in CTU tasks. MMFCTUB serves as a valuable resource for the research community, facilitating rigorous evaluation of MLLMs in the domain of CTU.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04643
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark
Yakun, Cui
Zhang, Yanting
Lei, Zhu
Xie, Jian
Kou, Zhizhuo
Du, Hang
Zhu, Zhenghao
Han, Sirui
Computational Engineering, Finance, and Science
The advent of multi-modal language models (MLLMs) has spurred research into their application across various table understanding tasks. However, their performance in credit table understanding (CTU) for financial credit review remains largely unexplored due to the following barriers: low data consistency, high annotation costs stemming from domain-specific knowledge and complex calculations, and evaluation paradigm gaps between benchmark and real-world scenarios. To address these challenges, we introduce MMFCTUB (Multi-Modal Financial Credit Table Understanding Benchmark), a practical benchmark, encompassing more than 7,600 high quality CTU samples across 5 table types. MMFCTUB employ a minimally supervised pipeline that adheres to inter-table constraints and maintains data distributions consistency. The benchmark leverages capacity-driven questions and mask-and-recovery strategy to evaluate models' cross-table structure perception, domain knowledge utilization, and numerical calculation capabilities. Utilizing MMFCTUB, we conduct comprehensive evaluations of both proprietary and open-source MLLMs, revealing their strengths and limitations in CTU tasks. MMFCTUB serves as a valuable resource for the research community, facilitating rigorous evaluation of MLLMs in the domain of CTU.
title MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2601.04643