Generating Realistic Tabular Data with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Dang, Gupta, Sunil, Do, Kien, Nguyen, Thin, Venkatesh, Svetha
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913566954618880
author Nguyen, Dang
Gupta, Sunil
Do, Kien
Nguyen, Thin
Venkatesh, Svetha
author_facet Nguyen, Dang
Gupta, Sunil
Do, Kien
Nguyen, Thin
Venkatesh, Svetha
contents While most generative models show achievements in image data generation, few are developed for tabular data generation. Recently, due to success of large language models (LLM) in diverse tasks, they have also been used for tabular data generation. However, these methods do not capture the correct correlation between the features and the target variable, hindering their applications in downstream predictive tasks. To address this problem, we propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data. First, we propose a novel permutation strategy for the input data in the fine-tuning phase. Second, we propose a feature-conditional sampling approach to generate synthetic samples. Finally, we generate the labels by constructing prompts based on the generated samples to query our fine-tuned LLM. Our extensive experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks. It also produces highly realistic synthetic samples in terms of quality and diversity. More importantly, classifiers trained with our synthetic data can even compete with classifiers trained with the original data on half of the benchmark datasets, which is a significant achievement in tabular data generation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21717
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generating Realistic Tabular Data with Large Language Models
Nguyen, Dang
Gupta, Sunil
Do, Kien
Nguyen, Thin
Venkatesh, Svetha
Machine Learning
Artificial Intelligence
While most generative models show achievements in image data generation, few are developed for tabular data generation. Recently, due to success of large language models (LLM) in diverse tasks, they have also been used for tabular data generation. However, these methods do not capture the correct correlation between the features and the target variable, hindering their applications in downstream predictive tasks. To address this problem, we propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data. First, we propose a novel permutation strategy for the input data in the fine-tuning phase. Second, we propose a feature-conditional sampling approach to generate synthetic samples. Finally, we generate the labels by constructing prompts based on the generated samples to query our fine-tuned LLM. Our extensive experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks. It also produces highly realistic synthetic samples in terms of quality and diversity. More importantly, classifiers trained with our synthetic data can even compete with classifiers trained with the original data on half of the benchmark datasets, which is a significant achievement in tabular data generation.
title Generating Realistic Tabular Data with Large Language Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.21717