Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xiaofeng, Ritter, Alan, Xu, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911086738931712
author Wu, Xiaofeng
Ritter, Alan
Xu, Wei
author_facet Wu, Xiaofeng
Ritter, Alan
Xu, Wei
contents Tables have gained significant attention in large language models (LLMs) and multimodal large language models (MLLMs) due to their complex and flexible structure. Unlike linear text inputs, tables are two-dimensional, encompassing formats that range from well-structured database tables to complex, multi-layered spreadsheets, each with different purposes. This diversity in format and purpose has led to the development of specialized methods and tasks, instead of universal approaches, making navigation of table understanding tasks challenging. To address these challenges, this paper introduces key concepts through a taxonomy of tabular input representations and an introduction of table understanding tasks. We highlight several critical gaps in the field that indicate the need for further research: (1) the predominance of retrieval-focused tasks that require minimal reasoning beyond mathematical and logical operations; (2) significant challenges faced by models when processing complex table structures, large-scale tables, length context, or multi-table scenarios; and (3) the limited generalization of models across different tabular representations and formats.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
Wu, Xiaofeng
Ritter, Alan
Xu, Wei
Computation and Language
Databases
Machine Learning
Tables have gained significant attention in large language models (LLMs) and multimodal large language models (MLLMs) due to their complex and flexible structure. Unlike linear text inputs, tables are two-dimensional, encompassing formats that range from well-structured database tables to complex, multi-layered spreadsheets, each with different purposes. This diversity in format and purpose has led to the development of specialized methods and tasks, instead of universal approaches, making navigation of table understanding tasks challenging. To address these challenges, this paper introduces key concepts through a taxonomy of tabular input representations and an introduction of table understanding tasks. We highlight several critical gaps in the field that indicate the need for further research: (1) the predominance of retrieval-focused tasks that require minimal reasoning beyond mathematical and logical operations; (2) significant challenges faced by models when processing complex table structures, large-scale tables, length context, or multi-table scenarios; and (3) the limited generalization of models across different tabular representations and formats.
title Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
topic Computation and Language
Databases
Machine Learning
url https://arxiv.org/abs/2508.00217