Does Training with Synthetic Data Truly Protect Privacy?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Yunpeng, Zhang, Jie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910834778701824
author Zhao, Yunpeng
Zhang, Jie
author_facet Zhao, Yunpeng
Zhang, Jie
contents As synthetic data becomes increasingly popular in machine learning tasks, numerous methods--without formal differential privacy guarantees--use synthetic data for training. These methods often claim, either explicitly or implicitly, to protect the privacy of the original training data. In this work, we explore four different training paradigms: coreset selection, dataset distillation, data-free knowledge distillation, and synthetic data generated from diffusion models. While all these methods utilize synthetic data for training, they lead to vastly different conclusions regarding privacy preservation. We caution that empirical approaches to preserving data privacy require careful and rigorous evaluation; otherwise, they risk providing a false sense of privacy.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12976
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Does Training with Synthetic Data Truly Protect Privacy?
Zhao, Yunpeng
Zhang, Jie
Cryptography and Security
Machine Learning
As synthetic data becomes increasingly popular in machine learning tasks, numerous methods--without formal differential privacy guarantees--use synthetic data for training. These methods often claim, either explicitly or implicitly, to protect the privacy of the original training data. In this work, we explore four different training paradigms: coreset selection, dataset distillation, data-free knowledge distillation, and synthetic data generated from diffusion models. While all these methods utilize synthetic data for training, they lead to vastly different conclusions regarding privacy preservation. We caution that empirical approaches to preserving data privacy require careful and rigorous evaluation; otherwise, they risk providing a false sense of privacy.
title Does Training with Synthetic Data Truly Protect Privacy?
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2502.12976