Quality Assessment of Tabular Data using Large Language Models and Code Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909798238257152 |
|---|---|
| author | Akella, Ashlesha Kaul, Akshar Narayanam, Krishnasuri Mehta, Sameep |
| author_facet | Akella, Ashlesha Kaul, Akshar Narayanam, Krishnasuri Mehta, Sameep |
| contents | Reliable data quality is crucial for downstream analysis of tabular datasets, yet rule-based validation often struggles with inefficiency, human intervention, and high computational costs. We present a three-stage framework that combines statistical inliner detection with LLM-driven rule and code generation. After filtering data samples through traditional clustering, we iteratively prompt LLMs to produce semantically valid quality rules and synthesize their executable validators through code-generating LLMs. To generate reliable quality rules, we aid LLMs with retrieval-augmented generation (RAG) by leveraging external knowledge sources and domain-specific few-shot examples. Robust guardrails ensure the accuracy and consistency of both rules and code snippets. Extensive evaluations on benchmark datasets confirm the effectiveness of our approach. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_10572 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Quality Assessment of Tabular Data using Large Language Models and Code Generation Akella, Ashlesha Kaul, Akshar Narayanam, Krishnasuri Mehta, Sameep Software Engineering Artificial Intelligence Databases Reliable data quality is crucial for downstream analysis of tabular datasets, yet rule-based validation often struggles with inefficiency, human intervention, and high computational costs. We present a three-stage framework that combines statistical inliner detection with LLM-driven rule and code generation. After filtering data samples through traditional clustering, we iteratively prompt LLMs to produce semantically valid quality rules and synthesize their executable validators through code-generating LLMs. To generate reliable quality rules, we aid LLMs with retrieval-augmented generation (RAG) by leveraging external knowledge sources and domain-specific few-shot examples. Robust guardrails ensure the accuracy and consistency of both rules and code snippets. Extensive evaluations on benchmark datasets confirm the effectiveness of our approach. |
| title | Quality Assessment of Tabular Data using Large Language Models and Code Generation |
| topic | Software Engineering Artificial Intelligence Databases |
| url | https://arxiv.org/abs/2509.10572 |