TabDiff: a Mixed-type Diffusion Model for Tabular Data Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shi, Juntong, Xu, Minkai, Hua, Harper, Zhang, Hengrui, Ermon, Stefano, Leskovec, Jure
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916616398176256
author Shi, Juntong
Xu, Minkai
Hua, Harper
Zhang, Hengrui
Ermon, Stefano
Leskovec, Jure
author_facet Shi, Juntong
Xu, Minkai
Hua, Harper
Zhang, Hengrui
Ermon, Stefano
Leskovec, Jure
contents Synthesizing high-quality tabular data is an important topic in many data science tasks, ranging from dataset augmentation to privacy protection. However, developing expressive generative models for tabular data is challenging due to its inherent heterogeneous data types, complex inter-correlations, and intricate column-wise distributions. In this paper, we introduce TabDiff, a joint diffusion framework that models all mixed-type distributions of tabular data in one model. Our key innovation is the development of a joint continuous-time diffusion process for numerical and categorical data, where we propose feature-wise learnable diffusion processes to counter the high disparity of different feature distributions. TabDiff is parameterized by a transformer handling different input types, and the entire framework can be efficiently optimized in an end-to-end fashion. We further introduce a mixed-type stochastic sampler to automatically correct the accumulated decoding error during sampling, and propose classifier-free guidance for conditional missing column value imputation. Comprehensive experiments on seven datasets demonstrate that TabDiff achieves superior average performance over existing competitive baselines across all eight metrics, with up to $22.5\%$ improvement over the state-of-the-art model on pair-wise column correlation estimations. Code is available at https://github.com/MinkaiXu/TabDiff.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20626
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TabDiff: a Mixed-type Diffusion Model for Tabular Data Generation
Shi, Juntong
Xu, Minkai
Hua, Harper
Zhang, Hengrui
Ermon, Stefano
Leskovec, Jure
Machine Learning
Synthesizing high-quality tabular data is an important topic in many data science tasks, ranging from dataset augmentation to privacy protection. However, developing expressive generative models for tabular data is challenging due to its inherent heterogeneous data types, complex inter-correlations, and intricate column-wise distributions. In this paper, we introduce TabDiff, a joint diffusion framework that models all mixed-type distributions of tabular data in one model. Our key innovation is the development of a joint continuous-time diffusion process for numerical and categorical data, where we propose feature-wise learnable diffusion processes to counter the high disparity of different feature distributions. TabDiff is parameterized by a transformer handling different input types, and the entire framework can be efficiently optimized in an end-to-end fashion. We further introduce a mixed-type stochastic sampler to automatically correct the accumulated decoding error during sampling, and propose classifier-free guidance for conditional missing column value imputation. Comprehensive experiments on seven datasets demonstrate that TabDiff achieves superior average performance over existing competitive baselines across all eight metrics, with up to $22.5\%$ improvement over the state-of-the-art model on pair-wise column correlation estimations. Code is available at https://github.com/MinkaiXu/TabDiff.
title TabDiff: a Mixed-type Diffusion Model for Tabular Data Generation
topic Machine Learning
url https://arxiv.org/abs/2410.20626