Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Borisova, Ekaterina, Barth, Fabio, Feldhus, Nils, Ahmad, Raia Abu, Ostendorff, Malte, Suarez, Pedro Ortiz, Rehm, Georg, Möller, Sebastian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911121353474048
author Borisova, Ekaterina
Barth, Fabio
Feldhus, Nils
Ahmad, Raia Abu
Ostendorff, Malte
Suarez, Pedro Ortiz
Rehm, Georg
Möller, Sebastian
author_facet Borisova, Ekaterina
Barth, Fabio
Feldhus, Nils
Ahmad, Raia Abu
Ostendorff, Malte
Suarez, Pedro Ortiz
Rehm, Georg
Möller, Sebastian
contents Tables are among the most widely used tools for representing structured data in research, business, medicine, and education. Although LLMs demonstrate strong performance in downstream tasks, their efficiency in processing tabular data remains underexplored. In this paper, we investigate the effectiveness of both text-based and multimodal LLMs on table understanding tasks through a cross-domain and cross-modality evaluation. Specifically, we compare their performance on tables from scientific vs. non-scientific contexts and examine their robustness on tables represented as images vs. text. Additionally, we conduct an interpretability analysis to measure context usage and input relevance. We also introduce the TableEval benchmark, comprising 3017 tables from scholarly publications, Wikipedia, and financial reports, where each table is provided in five different formats: Image, Dictionary, HTML, XML, and LaTeX. Our findings indicate that while LLMs maintain robustness across table modalities, they face significant challenges when processing scientific tables.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00152
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
Borisova, Ekaterina
Barth, Fabio
Feldhus, Nils
Ahmad, Raia Abu
Ostendorff, Malte
Suarez, Pedro Ortiz
Rehm, Georg
Möller, Sebastian
Computation and Language
Tables are among the most widely used tools for representing structured data in research, business, medicine, and education. Although LLMs demonstrate strong performance in downstream tasks, their efficiency in processing tabular data remains underexplored. In this paper, we investigate the effectiveness of both text-based and multimodal LLMs on table understanding tasks through a cross-domain and cross-modality evaluation. Specifically, we compare their performance on tables from scientific vs. non-scientific contexts and examine their robustness on tables represented as images vs. text. Additionally, we conduct an interpretability analysis to measure context usage and input relevance. We also introduce the TableEval benchmark, comprising 3017 tables from scholarly publications, Wikipedia, and financial reports, where each table is provided in five different formats: Image, Dictionary, HTML, XML, and LaTeX. Our findings indicate that while LLMs maintain robustness across table modalities, they face significant challenges when processing scientific tables.
title Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
topic Computation and Language
url https://arxiv.org/abs/2507.00152