Quality Assessment of Tabular Data using Large Language Models and Code Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Akella, Ashlesha, Kaul, Akshar, Narayanam, Krishnasuri, Mehta, Sameep
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909798238257152
author Akella, Ashlesha
Kaul, Akshar
Narayanam, Krishnasuri
Mehta, Sameep
author_facet Akella, Ashlesha
Kaul, Akshar
Narayanam, Krishnasuri
Mehta, Sameep
contents Reliable data quality is crucial for downstream analysis of tabular datasets, yet rule-based validation often struggles with inefficiency, human intervention, and high computational costs. We present a three-stage framework that combines statistical inliner detection with LLM-driven rule and code generation. After filtering data samples through traditional clustering, we iteratively prompt LLMs to produce semantically valid quality rules and synthesize their executable validators through code-generating LLMs. To generate reliable quality rules, we aid LLMs with retrieval-augmented generation (RAG) by leveraging external knowledge sources and domain-specific few-shot examples. Robust guardrails ensure the accuracy and consistency of both rules and code snippets. Extensive evaluations on benchmark datasets confirm the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quality Assessment of Tabular Data using Large Language Models and Code Generation
Akella, Ashlesha
Kaul, Akshar
Narayanam, Krishnasuri
Mehta, Sameep
Software Engineering
Artificial Intelligence
Databases
Reliable data quality is crucial for downstream analysis of tabular datasets, yet rule-based validation often struggles with inefficiency, human intervention, and high computational costs. We present a three-stage framework that combines statistical inliner detection with LLM-driven rule and code generation. After filtering data samples through traditional clustering, we iteratively prompt LLMs to produce semantically valid quality rules and synthesize their executable validators through code-generating LLMs. To generate reliable quality rules, we aid LLMs with retrieval-augmented generation (RAG) by leveraging external knowledge sources and domain-specific few-shot examples. Robust guardrails ensure the accuracy and consistency of both rules and code snippets. Extensive evaluations on benchmark datasets confirm the effectiveness of our approach.
title Quality Assessment of Tabular Data using Large Language Models and Code Generation
topic Software Engineering
Artificial Intelligence
Databases
url https://arxiv.org/abs/2509.10572