Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Wei, Zhou, Jun, Wang, Haoyu, Li, Zhenghao, He, Qikang, Han, Shaokun, Li, Guoliang, Zhou, Xuanhe, He, Yeye, Liu, Chunwei, Tang, Zirui, Wang, Bin, Tang, Shen, Zuo, Kai, Luo, Yuyu, Zheng, Zhenzhe, He, Conghui, Zhou, Jingren, Wu, Fan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914276768219136
author Zhou, Wei
Zhou, Jun
Wang, Haoyu
Li, Zhenghao
He, Qikang
Han, Shaokun
Li, Guoliang
Zhou, Xuanhe
He, Yeye
Liu, Chunwei
Tang, Zirui
Wang, Bin
Tang, Shen
Zuo, Kai
Luo, Yuyu
Zheng, Zhenzhe
He, Conghui
Zhou, Jingren
Wu, Fan
author_facet Zhou, Wei
Zhou, Jun
Wang, Haoyu
Li, Zhenghao
He, Qikang
Han, Shaokun
Li, Guoliang
Zhou, Xuanhe
He, Yeye
Liu, Chunwei
Tang, Zirui
Wang, Bin
Tang, Shen
Zuo, Kai
Luo, Yuyu
Zheng, Zhenzhe
He, Conghui
Zhou, Jingren
Wu, Fan
contents Data preparation aims to denoise raw datasets, uncover cross-dataset relationships, and extract valuable insights from them, which is essential for a wide range of data-centric applications. Driven by (i) rising demands for application-ready data (e.g., for analytics, visualization, decision-making), (ii) increasingly powerful LLM techniques, and (iii) the emergence of infrastructures that facilitate flexible agent construction (e.g., using Databricks Unity Catalog), LLM-enhanced methods are rapidly becoming a transformative and potentially dominant paradigm for data preparation. By investigating hundreds of recent literature works, this paper presents a systematic review of this evolving landscape, focusing on the use of LLM techniques to prepare data for diverse downstream tasks. First, we characterize the fundamental paradigm shift, from rule-based, model-specific pipelines to prompt-driven, context-aware, and agentic preparation workflows. Next, we introduce a task-centric taxonomy that organizes the field into three major tasks: data cleaning (e.g., standardization, error processing, imputation), data integration (e.g., entity matching, schema matching), and data enrichment (e.g., data annotation, profiling). For each task, we survey representative techniques, and highlight their respective strengths (e.g., improved generalization, semantic understanding) and limitations (e.g., the prohibitive cost of scaling LLMs, persistent hallucinations even in advanced agents, the mismatch between advanced methods and weak evaluation). Moreover, we analyze commonly used datasets and evaluation metrics (the empirical part). Finally, we discuss open research challenges and outline a forward-looking roadmap that emphasizes scalable LLM-data systems, principled designs for reliable agentic workflows, and robust evaluation protocols.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17058
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
Zhou, Wei
Zhou, Jun
Wang, Haoyu
Li, Zhenghao
He, Qikang
Han, Shaokun
Li, Guoliang
Zhou, Xuanhe
He, Yeye
Liu, Chunwei
Tang, Zirui
Wang, Bin
Tang, Shen
Zuo, Kai
Luo, Yuyu
Zheng, Zhenzhe
He, Conghui
Zhou, Jingren
Wu, Fan
Databases
Artificial Intelligence
Computation and Language
Machine Learning
Data preparation aims to denoise raw datasets, uncover cross-dataset relationships, and extract valuable insights from them, which is essential for a wide range of data-centric applications. Driven by (i) rising demands for application-ready data (e.g., for analytics, visualization, decision-making), (ii) increasingly powerful LLM techniques, and (iii) the emergence of infrastructures that facilitate flexible agent construction (e.g., using Databricks Unity Catalog), LLM-enhanced methods are rapidly becoming a transformative and potentially dominant paradigm for data preparation. By investigating hundreds of recent literature works, this paper presents a systematic review of this evolving landscape, focusing on the use of LLM techniques to prepare data for diverse downstream tasks. First, we characterize the fundamental paradigm shift, from rule-based, model-specific pipelines to prompt-driven, context-aware, and agentic preparation workflows. Next, we introduce a task-centric taxonomy that organizes the field into three major tasks: data cleaning (e.g., standardization, error processing, imputation), data integration (e.g., entity matching, schema matching), and data enrichment (e.g., data annotation, profiling). For each task, we survey representative techniques, and highlight their respective strengths (e.g., improved generalization, semantic understanding) and limitations (e.g., the prohibitive cost of scaling LLMs, persistent hallucinations even in advanced agents, the mismatch between advanced methods and weak evaluation). Moreover, we analyze commonly used datasets and evaluation metrics (the empirical part). Finally, we discuss open research challenges and outline a forward-looking roadmap that emphasizes scalable LLM-data systems, principled designs for reliable agentic workflows, and robust evaluation protocols.
title Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
topic Databases
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.17058