Benchmarking Differentially Private Tabular Data Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Kai, Li, Xiaochen, Gong, Chen, McKenna, Ryan, Wang, Tianhao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908659868499968
author Chen, Kai
Li, Xiaochen
Gong, Chen
McKenna, Ryan
Wang, Tianhao
author_facet Chen, Kai
Li, Xiaochen
Gong, Chen
McKenna, Ryan
Wang, Tianhao
contents Differentially private (DP) tabular data synthesis generates artificial data that preserves the statistical properties of private data while safeguarding individual privacy. The emergence of diverse algorithms in recent years has introduced challenges in practical applications, such as inconsistent data processing methods, the lack of in-depth algorithm analysis, and incomplete comparisons due to overlapping development timelines. These factors create significant obstacles to selecting appropriate algorithms. In this paper, we address these challenges by proposing a benchmark for evaluating tabular data synthesis methods. We present a unified evaluation framework that integrates data preprocessing, feature selection, and synthesis modules, facilitating fair and comprehensive comparisons. Our evaluation reveals that a significant utility-efficiency trade-off exists among current state-of-the-art methods. Some statistical methods are superior in synthesis utility, but their efficiency is not as good as most deep learning-based methods. Furthermore, we conduct an in-depth analysis of each module with experimental validation, offering theoretical insights into the strengths and limitations of different strategies. Our code is open-sourced via the link.\footnote{https://github.com/KaiChen9909/tab_bench}
format Preprint
id arxiv_https___arxiv_org_abs_2504_14061
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Differentially Private Tabular Data Synthesis
Chen, Kai
Li, Xiaochen
Gong, Chen
McKenna, Ryan
Wang, Tianhao
Cryptography and Security
Differentially private (DP) tabular data synthesis generates artificial data that preserves the statistical properties of private data while safeguarding individual privacy. The emergence of diverse algorithms in recent years has introduced challenges in practical applications, such as inconsistent data processing methods, the lack of in-depth algorithm analysis, and incomplete comparisons due to overlapping development timelines. These factors create significant obstacles to selecting appropriate algorithms. In this paper, we address these challenges by proposing a benchmark for evaluating tabular data synthesis methods. We present a unified evaluation framework that integrates data preprocessing, feature selection, and synthesis modules, facilitating fair and comprehensive comparisons. Our evaluation reveals that a significant utility-efficiency trade-off exists among current state-of-the-art methods. Some statistical methods are superior in synthesis utility, but their efficiency is not as good as most deep learning-based methods. Furthermore, we conduct an in-depth analysis of each module with experimental validation, offering theoretical insights into the strengths and limitations of different strategies. Our code is open-sourced via the link.\footnote{https://github.com/KaiChen9909/tab_bench}
title Benchmarking Differentially Private Tabular Data Synthesis
topic Cryptography and Security
url https://arxiv.org/abs/2504.14061