TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zheyuan, Shang, Liqiang, Chen, Junjie, Yang, Xun, Xu, Chenglong, Yuan, Bo, Jiao, Chenyuan, Sun, Yaoru, Zhao, Yilun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911656768962560
author Yang, Zheyuan
Shang, Liqiang
Chen, Junjie
Yang, Xun
Xu, Chenglong
Yuan, Bo
Jiao, Chenyuan
Sun, Yaoru
Zhao, Yilun
author_facet Yang, Zheyuan
Shang, Liqiang
Chen, Junjie
Yang, Xun
Xu, Chenglong
Yuan, Bo
Jiao, Chenyuan
Sun, Yaoru
Zhao, Yilun
contents We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each instance is expanded into 10 distinct visual variants through our multi-style rendering and transformation pipeline. This process encompasses diverse scenario styles, robustness perturbations, and vision-only configurations, culminating in 30,000 multimodal samples for a multi-dimensional evaluation. We conduct an extensive evaluation of 29 state-of-the-art open-source and proprietary foundation models on TableVista. Through comprehensive quantitative and qualitative analysis, we find that while evaluated models remain largely stable across diverse rendering styles, they exhibit pronounced performance degradation on complex structural layouts and vision-only settings, revealing that current models struggle to maintain reasoning consistency when structural complexity combines with visually integrated presentations. These findings highlight critical gaps in current multimodal capabilities, providing insights for advancing more robust and reliable table understanding models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05955
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
Yang, Zheyuan
Shang, Liqiang
Chen, Junjie
Yang, Xun
Xu, Chenglong
Yuan, Bo
Jiao, Chenyuan
Sun, Yaoru
Zhao, Yilun
Computation and Language
Computer Vision and Pattern Recognition
We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each instance is expanded into 10 distinct visual variants through our multi-style rendering and transformation pipeline. This process encompasses diverse scenario styles, robustness perturbations, and vision-only configurations, culminating in 30,000 multimodal samples for a multi-dimensional evaluation. We conduct an extensive evaluation of 29 state-of-the-art open-source and proprietary foundation models on TableVista. Through comprehensive quantitative and qualitative analysis, we find that while evaluated models remain largely stable across diverse rendering styles, they exhibit pronounced performance degradation on complex structural layouts and vision-only settings, revealing that current models struggle to maintain reasoning consistency when structural complexity combines with visually integrated presentations. These findings highlight critical gaps in current multimodal capabilities, providing insights for advancing more robust and reliable table understanding models.
title TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.05955