FASTGEN: Fast and Cost-Effective Synthetic Tabular Data Generation with LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nguyen, Anh, Schafft, Sam, Hale, Nicholas, Alfaro, John
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909698031091712
author Nguyen, Anh
Schafft, Sam
Hale, Nicholas
Alfaro, John
author_facet Nguyen, Anh
Schafft, Sam
Hale, Nicholas
Alfaro, John
contents Synthetic data generation has emerged as an invaluable solution in scenarios where real-world data collection and usage are limited by cost and scarcity. Large language models (LLMs) have demonstrated remarkable capabilities in producing high-fidelity, domain-relevant samples across various fields. However, existing approaches that directly use LLMs to generate each record individually impose prohibitive time and cost burdens, particularly when large volumes of synthetic data are required. In this work, we propose a fast, cost-effective method for realistic tabular data synthesis that leverages LLMs to infer and encode each field's distribution into a reusable sampling script. By automatically classifying fields into numerical, categorical, or free-text types, the LLM generates distribution-based scripts that can efficiently produce diverse, realistic datasets at scale without continuous model inference. Experimental results show that our approach outperforms traditional direct methods in both diversity and data realism, substantially reducing the burden of high-volume synthetic data generation. We plan to apply this methodology to accelerate testing in production pipelines, thereby shortening development cycles and improving overall system efficiency. We believe our insights and lessons learned will aid researchers and practitioners seeking scalable, cost-effective solutions for synthetic data generation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FASTGEN: Fast and Cost-Effective Synthetic Tabular Data Generation with LLMs
Nguyen, Anh
Schafft, Sam
Hale, Nicholas
Alfaro, John
Machine Learning
Artificial Intelligence
Synthetic data generation has emerged as an invaluable solution in scenarios where real-world data collection and usage are limited by cost and scarcity. Large language models (LLMs) have demonstrated remarkable capabilities in producing high-fidelity, domain-relevant samples across various fields. However, existing approaches that directly use LLMs to generate each record individually impose prohibitive time and cost burdens, particularly when large volumes of synthetic data are required. In this work, we propose a fast, cost-effective method for realistic tabular data synthesis that leverages LLMs to infer and encode each field's distribution into a reusable sampling script. By automatically classifying fields into numerical, categorical, or free-text types, the LLM generates distribution-based scripts that can efficiently produce diverse, realistic datasets at scale without continuous model inference. Experimental results show that our approach outperforms traditional direct methods in both diversity and data realism, substantially reducing the burden of high-volume synthetic data generation. We plan to apply this methodology to accelerate testing in production pipelines, thereby shortening development cycles and improving overall system efficiency. We believe our insights and lessons learned will aid researchers and practitioners seeking scalable, cost-effective solutions for synthetic data generation.
title FASTGEN: Fast and Cost-Effective Synthetic Tabular Data Generation with LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2507.15839