Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Wei, Zhou, Jun, Wang, Haoyu, Li, Zhenghao, He, Qikang, Han, Shaokun, Li, Guoliang, Zhou, Xuanhe, He, Yeye, Liu, Chunwei, Tang, Zirui, Wang, Bin, Tang, Shen, Zuo, Kai, Luo, Yuyu, Zheng, Zhenzhe, He, Conghui, Zhou, Jingren, Wu, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automating Database-Native Function Code Synthesis with LLMs
by: Zhou, Wei, et al.
Published: (2026)
by: Zhou, Wei, et al.
Published: (2026)
MoDora: Tree-Based Semi-Structured Document Analysis System
by: Xu, Bangrui, et al.
Published: (2026)
by: Xu, Bangrui, et al.
Published: (2026)
LLM/Agent-as-Data-Analyst: A Survey
by: Tang, Zirui, et al.
Published: (2025)
by: Tang, Zirui, et al.
Published: (2025)
PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation
by: Zhou, Wei, et al.
Published: (2025)
by: Zhou, Wei, et al.
Published: (2025)
A Survey of LLM $\times$ DATA
by: Zhou, Xuanhe, et al.
Published: (2025)
by: Zhou, Xuanhe, et al.
Published: (2025)
Qute: Towards Quantum-Native Database
by: Chen, Muzhi, et al.
Published: (2026)
by: Chen, Muzhi, et al.
Published: (2026)
FeatInsight: An Online ML Feature Management System on 4Paradigm Sage-Studio Platform
by: Tong, Xin, et al.
Published: (2025)
by: Tong, Xin, et al.
Published: (2025)
ST-Raptor: LLM-Powered Semi-Structured Table Question Answering
by: Tang, Zirui, et al.
Published: (2025)
by: Tang, Zirui, et al.
Published: (2025)
LLM-Enhanced Data Management
by: Zhou, Xuanhe, et al.
Published: (2024)
by: Zhou, Xuanhe, et al.
Published: (2024)
CrackSQL: A Hybrid SQL Dialect Translation System Powered by Large Language Models
by: Zhou, Wei, et al.
Published: (2025)
by: Zhou, Wei, et al.
Published: (2025)
The Dawn of Natural Language to SQL: Are We Fully Ready?
by: Li, Boyan, et al.
Published: (2024)
by: Li, Boyan, et al.
Published: (2024)
ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
by: Tou, Huaixiao, et al.
Published: (2025)
by: Tou, Huaixiao, et al.
Published: (2025)
ST-Raptor: An Agentic System for Semi-Structured Table QA
by: Qu, Jinxiu, et al.
Published: (2026)
by: Qu, Jinxiu, et al.
Published: (2026)
Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)
by: Tang, Zichen, et al.
Published: (2026)
by: Tang, Zichen, et al.
Published: (2026)
Generating Robotic Control Strategies With LLMs : Via Human–Robot Voice Interaction
by: Xiaopeng Wang, et al.
Published: (2025)
by: Xiaopeng Wang, et al.
Published: (2025)
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
by: Xu, Bangrui, et al.
Published: (2026)
by: Xu, Bangrui, et al.
Published: (2026)
Generalized Category Discovery in Event-Centric Contexts: Latent Pattern Mining with LLMs
by: Luo, Yi, et al.
Published: (2025)
by: Luo, Yi, et al.
Published: (2025)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Clean Up the Mess: Addressing Data Pollution in Cryptocurrency Abuse Reporting Services
by: Gomez, Gibran, et al.
Published: (2024)
by: Gomez, Gibran, et al.
Published: (2024)
A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
"Get Ready for Migration: Clean-Up Your Collection"--What Does that Mean?
by: Lawrence, Peg, et al.
Published: (2008)
by: Lawrence, Peg, et al.
Published: (2008)
GaussMaster: An LLM-based Database Copilot System
by: Zhou, Wei, et al.
Published: (2025)
by: Zhou, Wei, et al.
Published: (2025)
Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs
by: Zhuang, Kai, et al.
Published: (2025)
by: Zhuang, Kai, et al.
Published: (2025)
Mono: Is Your "Clean" Vulnerability Dataset Really Solvable? Exposing and Trapping Undecidable Patches and Beyond
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
Dial: A Knowledge-Grounded Dialect-Specific NL2SQL System
by: Zhang, Xiang, et al.
Published: (2026)
by: Zhang, Xiang, et al.
Published: (2026)
Similarity-Based Domain Adaptation with LLMs
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
PromptKeeper: Safeguarding System Prompts for LLMs
by: Jiang, Zhifeng, et al.
Published: (2024)
by: Jiang, Zhifeng, et al.
Published: (2024)
LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
by: He, Ziyuan, et al.
Published: (2025)
by: He, Ziyuan, et al.
Published: (2025)
TableLoRA: Low-rank Adaptation on Table Structure Understanding for Large Language Models
by: He, Xinyi, et al.
Published: (2025)
by: He, Xinyi, et al.
Published: (2025)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
by: Sun, He, et al.
Published: (2026)
by: Sun, He, et al.
Published: (2026)
Revolutionizing Database Q&A with Large Language Models: Comprehensive Benchmark and Evaluation
by: Zheng, Yihang, et al.
Published: (2024)
by: Zheng, Yihang, et al.
Published: (2024)
A Retrieval-Augmented Knowledge Mining Method with Deep Thinking LLMs for Biomedical Research and Clinical Support
by: Feng, Yichun, et al.
Published: (2025)
by: Feng, Yichun, et al.
Published: (2025)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
by: Deng, Jinhong, et al.
Published: (2025)
by: Deng, Jinhong, et al.
Published: (2025)
Short-Path Prompting in LLMs: Analyzing Reasoning Instability and Solutions for Robust Performance
by: Tang, Zuoli, et al.
Published: (2025)
by: Tang, Zuoli, et al.
Published: (2025)
OpenMLDB: A Real-Time Relational Data Feature Computation System for Online ML
by: Zhou, Xuanhe, et al.
Published: (2025)
by: Zhou, Xuanhe, et al.
Published: (2025)
Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Generator-Validator Fine-tuning
by: Xing, Junjie, et al.
Published: (2024)
by: Xing, Junjie, et al.
Published: (2024)
Constraining the Galactic bar and spiral pattern speeds with the Hyades tidal stream
by: Zhou, Zi-yi, et al.
Published: (2026)
by: Zhou, Zi-yi, et al.
Published: (2026)
Will LLMs be Professional at Fund Investment? DeepFund: A Live Arena Perspective
by: Li, Changlun, et al.
Published: (2025)
by: Li, Changlun, et al.
Published: (2025)
Affine Dependence of Network Controllability/Observability on Its Subsystem Parameters and Connections
by: Zhou, Tong, et al.
Published: (2019)
by: Zhou, Tong, et al.
Published: (2019)
Similar Items
-
Automating Database-Native Function Code Synthesis with LLMs
by: Zhou, Wei, et al.
Published: (2026) -
MoDora: Tree-Based Semi-Structured Document Analysis System
by: Xu, Bangrui, et al.
Published: (2026) -
LLM/Agent-as-Data-Analyst: A Survey
by: Tang, Zirui, et al.
Published: (2025) -
PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation
by: Zhou, Wei, et al.
Published: (2025) -
A Survey of LLM $\times$ DATA
by: Zhou, Xuanhe, et al.
Published: (2025)