Differentially Private Synthetic Data Generation for Relational Databases

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Alimohammadi, Kaveh, Wang, Hao, Gulati, Ojas, Srivastava, Akash, Azizan, Navid
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915108682203136
author Alimohammadi, Kaveh
Wang, Hao
Gulati, Ojas
Srivastava, Akash
Azizan, Navid
author_facet Alimohammadi, Kaveh
Wang, Hao
Gulati, Ojas
Srivastava, Akash
Azizan, Navid
contents Existing differentially private (DP) synthetic data generation mechanisms typically assume a single-source table. In practice, data is often distributed across multiple tables with relationships across tables. In this paper, we introduce the first-of-its-kind algorithm that can be combined with any existing DP mechanisms to generate synthetic relational databases. Our algorithm iteratively refines the relationship between individual synthetic tables to minimize their approximation errors in terms of low-order marginal distributions while maintaining referential integrity. This algorithm eliminates the need to flatten a relational database into a master table (saving space), operates efficiently (saving time), and scales effectively to high-dimensional data. We provide both DP and theoretical utility guarantees for our algorithm. Through numerical experiments on real-world datasets, we demonstrate the effectiveness of our method in preserving fidelity to the original data.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18670
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Differentially Private Synthetic Data Generation for Relational Databases
Alimohammadi, Kaveh
Wang, Hao
Gulati, Ojas
Srivastava, Akash
Azizan, Navid
Machine Learning
Cryptography and Security
Databases
Existing differentially private (DP) synthetic data generation mechanisms typically assume a single-source table. In practice, data is often distributed across multiple tables with relationships across tables. In this paper, we introduce the first-of-its-kind algorithm that can be combined with any existing DP mechanisms to generate synthetic relational databases. Our algorithm iteratively refines the relationship between individual synthetic tables to minimize their approximation errors in terms of low-order marginal distributions while maintaining referential integrity. This algorithm eliminates the need to flatten a relational database into a master table (saving space), operates efficiently (saving time), and scales effectively to high-dimensional data. We provide both DP and theoretical utility guarantees for our algorithm. Through numerical experiments on real-world datasets, we demonstrate the effectiveness of our method in preserving fidelity to the original data.
title Differentially Private Synthetic Data Generation for Relational Databases
topic Machine Learning
Cryptography and Security
Databases
url https://arxiv.org/abs/2405.18670