SMOTE-DP: Improving Privacy-Utility Tradeoff with Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yan, Malin, Bradley, Kantarcioglu, Murat
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908389723865088
author Zhou, Yan
Malin, Bradley
Kantarcioglu, Murat
author_facet Zhou, Yan
Malin, Bradley
Kantarcioglu, Murat
contents Privacy-preserving data publication, including synthetic data sharing, often experiences trade-offs between privacy and utility. Synthetic data is generally more effective than data anonymization in balancing this trade-off, however, not without its own challenges. Synthetic data produced by generative models trained on source data may inadvertently reveal information about outliers. Techniques specifically designed for preserving privacy, such as introducing noise to satisfy differential privacy, often incur unpredictable and significant losses in utility. In this work we show that, with the right mechanism of synthetic data generation, we can achieve strong privacy protection without significant utility loss. Synthetic data generators producing contracting data patterns, such as Synthetic Minority Over-sampling Technique (SMOTE), can enhance a differentially private data generator, leveraging the strengths of both. We prove in theory and through empirical demonstration that this SMOTE-DP technique can produce synthetic data that not only ensures robust privacy protection but maintains utility in downstream learning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01907
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SMOTE-DP: Improving Privacy-Utility Tradeoff with Synthetic Data
Zhou, Yan
Malin, Bradley
Kantarcioglu, Murat
Machine Learning
Cryptography and Security
Privacy-preserving data publication, including synthetic data sharing, often experiences trade-offs between privacy and utility. Synthetic data is generally more effective than data anonymization in balancing this trade-off, however, not without its own challenges. Synthetic data produced by generative models trained on source data may inadvertently reveal information about outliers. Techniques specifically designed for preserving privacy, such as introducing noise to satisfy differential privacy, often incur unpredictable and significant losses in utility. In this work we show that, with the right mechanism of synthetic data generation, we can achieve strong privacy protection without significant utility loss. Synthetic data generators producing contracting data patterns, such as Synthetic Minority Over-sampling Technique (SMOTE), can enhance a differentially private data generator, leveraging the strengths of both. We prove in theory and through empirical demonstration that this SMOTE-DP technique can produce synthetic data that not only ensures robust privacy protection but maintains utility in downstream learning tasks.
title SMOTE-DP: Improving Privacy-Utility Tradeoff with Synthetic Data
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2506.01907