Multimodal Tabular Reasoning with Privileged Structured Information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Jun-Peng, Xia, Yu, Sun, Hai-Long, Lu, Shiyin, Chen, Qing-Guo, Luo, Weihua, Zhang, Kaifu, Zhan, De-Chuan, Ye, Han-Jia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912413091102720
author Jiang, Jun-Peng
Xia, Yu
Sun, Hai-Long
Lu, Shiyin
Chen, Qing-Guo
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
author_facet Jiang, Jun-Peng
Xia, Yu
Sun, Hai-Long
Lu, Shiyin
Chen, Qing-Guo
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
contents Tabular reasoning involves multi-step information extraction and logical inference over tabular data. While recent advances have leveraged large language models (LLMs) for reasoning over structured tables, such high-quality textual representations are often unavailable in real-world settings, where tables typically appear as images. In this paper, we tackle the task of tabular reasoning from table images, leveraging privileged structured information available during training to enhance multimodal large language models (MLLMs). The key challenges lie in the complexity of accurately aligning structured information with visual representations, and in effectively transferring structured reasoning skills to MLLMs despite the input modality gap. To address these, we introduce TabUlar Reasoning with Bridged infOrmation ({\sc Turbo}), a new framework for multimodal tabular reasoning with privileged structured tables. {\sc Turbo} benefits from a structure-aware reasoning trace generator based on DeepSeek-R1, contributing to high-quality modality-bridged data. On this basis, {\sc Turbo} repeatedly generates and selects the advantageous reasoning paths, further enhancing the model's tabular reasoning ability. Experimental results demonstrate that, with limited ($9$k) data, {\sc Turbo} achieves state-of-the-art performance ($+7.2\%$ vs. previous SOTA) across multiple datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Tabular Reasoning with Privileged Structured Information
Jiang, Jun-Peng
Xia, Yu
Sun, Hai-Long
Lu, Shiyin
Chen, Qing-Guo
Luo, Weihua
Zhang, Kaifu
Zhan, De-Chuan
Ye, Han-Jia
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Tabular reasoning involves multi-step information extraction and logical inference over tabular data. While recent advances have leveraged large language models (LLMs) for reasoning over structured tables, such high-quality textual representations are often unavailable in real-world settings, where tables typically appear as images. In this paper, we tackle the task of tabular reasoning from table images, leveraging privileged structured information available during training to enhance multimodal large language models (MLLMs). The key challenges lie in the complexity of accurately aligning structured information with visual representations, and in effectively transferring structured reasoning skills to MLLMs despite the input modality gap. To address these, we introduce TabUlar Reasoning with Bridged infOrmation ({\sc Turbo}), a new framework for multimodal tabular reasoning with privileged structured tables. {\sc Turbo} benefits from a structure-aware reasoning trace generator based on DeepSeek-R1, contributing to high-quality modality-bridged data. On this basis, {\sc Turbo} repeatedly generates and selects the advantageous reasoning paths, further enhancing the model's tabular reasoning ability. Experimental results demonstrate that, with limited ($9$k) data, {\sc Turbo} achieves state-of-the-art performance ($+7.2\%$ vs. previous SOTA) across multiple datasets.
title Multimodal Tabular Reasoning with Privileged Structured Information
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.04088