TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911656768962560 |
|---|---|
| author | Yang, Zheyuan Shang, Liqiang Chen, Junjie Yang, Xun Xu, Chenglong Yuan, Bo Jiao, Chenyuan Sun, Yaoru Zhao, Yilun |
| author_facet | Yang, Zheyuan Shang, Liqiang Chen, Junjie Yang, Xun Xu, Chenglong Yuan, Bo Jiao, Chenyuan Sun, Yaoru Zhao, Yilun |
| contents | We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each instance is expanded into 10 distinct visual variants through our multi-style rendering and transformation pipeline. This process encompasses diverse scenario styles, robustness perturbations, and vision-only configurations, culminating in 30,000 multimodal samples for a multi-dimensional evaluation. We conduct an extensive evaluation of 29 state-of-the-art open-source and proprietary foundation models on TableVista. Through comprehensive quantitative and qualitative analysis, we find that while evaluated models remain largely stable across diverse rendering styles, they exhibit pronounced performance degradation on complex structural layouts and vision-only settings, revealing that current models struggle to maintain reasoning consistency when structural complexity combines with visually integrated presentations. These findings highlight critical gaps in current multimodal capabilities, providing insights for advancing more robust and reliable table understanding models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_05955 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity Yang, Zheyuan Shang, Liqiang Chen, Junjie Yang, Xun Xu, Chenglong Yuan, Bo Jiao, Chenyuan Sun, Yaoru Zhao, Yilun Computation and Language Computer Vision and Pattern Recognition We introduce TableVista, a comprehensive benchmark for evaluating foundation models in multimodal table reasoning under visual and structural complexity. TableVista consists of 3,000 high-quality table reasoning problems, where each instance is expanded into 10 distinct visual variants through our multi-style rendering and transformation pipeline. This process encompasses diverse scenario styles, robustness perturbations, and vision-only configurations, culminating in 30,000 multimodal samples for a multi-dimensional evaluation. We conduct an extensive evaluation of 29 state-of-the-art open-source and proprietary foundation models on TableVista. Through comprehensive quantitative and qualitative analysis, we find that while evaluated models remain largely stable across diverse rendering styles, they exhibit pronounced performance degradation on complex structural layouts and vision-only settings, revealing that current models struggle to maintain reasoning consistency when structural complexity combines with visually integrated presentations. These findings highlight critical gaps in current multimodal capabilities, providing insights for advancing more robust and reliable table understanding models. |
| title | TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.05955 |