TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Jongha, Bae, Minseong, Lee, Sanghyeok, Yoon, Jinsung, Kim, Hyunwoo J.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911270118096896
author Kim, Jongha
Bae, Minseong
Lee, Sanghyeok
Yoon, Jinsung
Kim, Hyunwoo J.
author_facet Kim, Jongha
Bae, Minseong
Lee, Sanghyeok
Yoon, Jinsung
Kim, Hyunwoo J.
contents Table images present unique challenges for effective and efficient understanding due to the need for question-specific focus and the presence of redundant background regions. Existing Multimodal Large Language Model (MLLM) approaches often overlook these characteristics, resulting in uninformative and redundant visual representations. To address these issues, we aim to generate visual features that are both informative and compact to improve table understanding. We first propose progressive question conditioning, which injects the question into Vision Transformer layers with gradually increasing frequency, considering each layer's capacity to handle additional information, to generate question-aware visual features. To reduce redundancy, we introduce a pruning strategy that discards background tokens, thereby improving efficiency. To mitigate information loss from pruning, we further propose token focusing, a training strategy that encourages the model to concentrate essential information in the retained tokens. By combining these approaches, we present TabFlash, an efficient and effective MLLM for table understanding. TabFlash achieves state-of-the-art performance, outperforming both open-source and proprietary MLLMs, while requiring 27% less FLOPs and 30% less memory usage compared to the second-best MLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13283
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
Kim, Jongha
Bae, Minseong
Lee, Sanghyeok
Yoon, Jinsung
Kim, Hyunwoo J.
Computer Vision and Pattern Recognition
Table images present unique challenges for effective and efficient understanding due to the need for question-specific focus and the presence of redundant background regions. Existing Multimodal Large Language Model (MLLM) approaches often overlook these characteristics, resulting in uninformative and redundant visual representations. To address these issues, we aim to generate visual features that are both informative and compact to improve table understanding. We first propose progressive question conditioning, which injects the question into Vision Transformer layers with gradually increasing frequency, considering each layer's capacity to handle additional information, to generate question-aware visual features. To reduce redundancy, we introduce a pruning strategy that discards background tokens, thereby improving efficiency. To mitigate information loss from pruning, we further propose token focusing, a training strategy that encourages the model to concentrate essential information in the retained tokens. By combining these approaches, we present TabFlash, an efficient and effective MLLM for table understanding. TabFlash achieves state-of-the-art performance, outperforming both open-source and proprietary MLLMs, while requiring 27% less FLOPs and 30% less memory usage compared to the second-best MLLM.
title TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.13283