Empowering Tabular Data Preparation with Language Models: Why and How?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Mengshi, Sun, Yuxiang, Li, Tengchao, Wang, Jianwei, Wang, Kai, Lin, Xuemin, Zhang, Ying, Zhang, Wenjie
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918112345980928
author Chen, Mengshi
Sun, Yuxiang
Li, Tengchao
Wang, Jianwei
Wang, Kai
Lin, Xuemin
Zhang, Ying
Zhang, Wenjie
author_facet Chen, Mengshi
Sun, Yuxiang
Li, Tengchao
Wang, Jianwei
Wang, Kai
Lin, Xuemin
Zhang, Ying
Zhang, Wenjie
contents Data preparation is a critical step in enhancing the usability of tabular data and thus boosts downstream data-driven tasks. Traditional methods often face challenges in capturing the intricate relationships within tables and adapting to the tasks involved. Recent advances in Language Models (LMs), especially in Large Language Models (LLMs), offer new opportunities to automate and support tabular data preparation. However, why LMs suit tabular data preparation (i.e., how their capabilities match task demands) and how to use them effectively across phases still remain to be systematically explored. In this survey, we systematically analyze the role of LMs in enhancing tabular data preparation processes, focusing on four core phases: data acquisition, integration, cleaning, and transformation. For each phase, we present an integrated analysis of how LMs can be combined with other components for different preparation tasks, highlight key advancements, and outline prospective pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Empowering Tabular Data Preparation with Language Models: Why and How?
Chen, Mengshi
Sun, Yuxiang
Li, Tengchao
Wang, Jianwei
Wang, Kai
Lin, Xuemin
Zhang, Ying
Zhang, Wenjie
Artificial Intelligence
68T50
I.2.7
Data preparation is a critical step in enhancing the usability of tabular data and thus boosts downstream data-driven tasks. Traditional methods often face challenges in capturing the intricate relationships within tables and adapting to the tasks involved. Recent advances in Language Models (LMs), especially in Large Language Models (LLMs), offer new opportunities to automate and support tabular data preparation. However, why LMs suit tabular data preparation (i.e., how their capabilities match task demands) and how to use them effectively across phases still remain to be systematically explored. In this survey, we systematically analyze the role of LMs in enhancing tabular data preparation processes, focusing on four core phases: data acquisition, integration, cleaning, and transformation. For each phase, we present an integrated analysis of how LMs can be combined with other components for different preparation tasks, highlight key advancements, and outline prospective pipelines.
title Empowering Tabular Data Preparation with Language Models: Why and How?
topic Artificial Intelligence
68T50
I.2.7
url https://arxiv.org/abs/2508.01556