DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fan, Meihao, Fan, Ju, Zhang, Yuxin, Zhang, Shaolei, Du, Xiaoyong, Song, Jie, Li, Peng, Jiang, Fuxin, Zhang, Tieying, Chen, Jianjun
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914313737863168
author Fan, Meihao
Fan, Ju
Zhang, Yuxin
Zhang, Shaolei
Du, Xiaoyong
Song, Jie
Li, Peng
Jiang, Fuxin
Zhang, Tieying
Chen, Jianjun
author_facet Fan, Meihao
Fan, Ju
Zhang, Yuxin
Zhang, Shaolei
Du, Xiaoyong
Song, Jie
Li, Peng
Jiang, Fuxin
Zhang, Tieying
Chen, Jianjun
contents Data preparation, which aims to transform heterogeneous and noisy raw tables into analysis-ready data, remains a major bottleneck in data science. Recent approaches leverage large language models (LLMs) to automate data preparation from natural language specifications. However, existing LLM-powered methods either make decisions without grounding in intermediate execution results, or rely on linear interaction processes that offer limited support for revising earlier decisions. To address these limitations, we propose DeepPrep, an LLM-powered agentic system for autonomous data preparation. DeepPrep constructs data preparation pipelines through iterative, execution-grounded interaction with an environment that materializes intermediate table states and returns runtime feedback. To overcome the limitations of linear interaction, DeepPrep organizes pipeline construction with tree-based agentic reasoning, enabling structured exploration and non-local revision based on execution feedback. To enable effective learning of such behaviors, we propose a progressive agentic training framework, together with data synthesis that supplies diverse and complex ADP tasks. Extensive experiments show that DeepPrep achieves data preparation accuracy comparable to strong closed-source models (e.g., GPT-5) while incurring 15x lower inference cost, while establishing state-of-the-art performance among open-source baselines and generalizing effectively across diverse datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07371
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
Fan, Meihao
Fan, Ju
Zhang, Yuxin
Zhang, Shaolei
Du, Xiaoyong
Song, Jie
Li, Peng
Jiang, Fuxin
Zhang, Tieying
Chen, Jianjun
Databases
Data preparation, which aims to transform heterogeneous and noisy raw tables into analysis-ready data, remains a major bottleneck in data science. Recent approaches leverage large language models (LLMs) to automate data preparation from natural language specifications. However, existing LLM-powered methods either make decisions without grounding in intermediate execution results, or rely on linear interaction processes that offer limited support for revising earlier decisions. To address these limitations, we propose DeepPrep, an LLM-powered agentic system for autonomous data preparation. DeepPrep constructs data preparation pipelines through iterative, execution-grounded interaction with an environment that materializes intermediate table states and returns runtime feedback. To overcome the limitations of linear interaction, DeepPrep organizes pipeline construction with tree-based agentic reasoning, enabling structured exploration and non-local revision based on execution feedback. To enable effective learning of such behaviors, we propose a progressive agentic training framework, together with data synthesis that supplies diverse and complex ADP tasks. Extensive experiments show that DeepPrep achieves data preparation accuracy comparable to strong closed-source models (e.g., GPT-5) while incurring 15x lower inference cost, while establishing state-of-the-art performance among open-source baselines and generalizing effectively across diverse datasets.
title DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
topic Databases
url https://arxiv.org/abs/2602.07371