ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Alqurnawi, Yahia, Biswas, Preetom, Rao, Anmol, Anvekar, Tejas, Baral, Chitta, Gupta, Vivek
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911452365848576
author Alqurnawi, Yahia
Biswas, Preetom
Rao, Anmol
Anvekar, Tejas
Baral, Chitta
Gupta, Vivek
author_facet Alqurnawi, Yahia
Biswas, Preetom
Rao, Anmol
Anvekar, Tejas
Baral, Chitta
Gupta, Vivek
contents Multimodal Large Language Models (mLLMs) are often used to answer questions in structured data such as tables in Markdown, JSON, and images. While these models can often give correct answers, users also need to know where those answers come from. In this work, we study structured data attribution/citation, which is the ability of the models to point to the specific rows and columns that support an answer. We evaluate several mLLMs across different table formats and prompting strategies. Our results show a clear gap between question answering and evidence attribution. Although question answering accuracy remains moderate, attribution accuracy is much lower, near random for JSON inputs, across all models. We also find that models are more reliable at citing rows than columns, and struggle more with textual formats than images. Finally, we observe notable differences across model families. Overall, our findings show that current mLLMs are unreliable at providing fine-grained, trustworthy attribution for structured data, which limits their usage in applications requiring transparency and traceability.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15769
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
Alqurnawi, Yahia
Biswas, Preetom
Rao, Anmol
Anvekar, Tejas
Baral, Chitta
Gupta, Vivek
Computation and Language
Multimodal Large Language Models (mLLMs) are often used to answer questions in structured data such as tables in Markdown, JSON, and images. While these models can often give correct answers, users also need to know where those answers come from. In this work, we study structured data attribution/citation, which is the ability of the models to point to the specific rows and columns that support an answer. We evaluate several mLLMs across different table formats and prompting strategies. Our results show a clear gap between question answering and evidence attribution. Although question answering accuracy remains moderate, attribution accuracy is much lower, near random for JSON inputs, across all models. We also find that models are more reliable at citing rows than columns, and struggle more with textual formats than images. Finally, we observe notable differences across model families. Overall, our findings show that current mLLMs are unreliable at providing fine-grained, trustworthy attribution for structured data, which limits their usage in applications requiring transparency and traceability.
title ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
topic Computation and Language
url https://arxiv.org/abs/2602.15769