Private Training & Data Generation by Clustering Embeddings

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Felix, Zhou, Samson, Mirrokni, Vahab, Epasto, Alessandro, Cohen-Addad, Vincent
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909654071640064
author Zhou, Felix
Zhou, Samson
Mirrokni, Vahab
Epasto, Alessandro
Cohen-Addad, Vincent
author_facet Zhou, Felix
Zhou, Samson
Mirrokni, Vahab
Epasto, Alessandro
Cohen-Addad, Vincent
contents Deep neural networks often use large, high-quality datasets to achieve high performance on many machine learning tasks. When training involves potentially sensitive data, this process can raise privacy concerns, as large models have been shown to unintentionally memorize and reveal sensitive information, including reconstructing entire training samples. Differential privacy (DP) provides a robust framework for protecting individual data and in particular, a new approach to privately training deep neural networks is to approximate the input dataset with a privately generated synthetic dataset, before any subsequent training algorithm. We introduce a novel principled method for DP synthetic image embedding generation, based on fitting a Gaussian Mixture Model (GMM) in an appropriate embedding space using DP clustering. Our method provably learns a GMM under separation conditions. Empirically, a simple two-layer neural network trained on synthetically generated embeddings achieves state-of-the-art (SOTA) classification accuracy on standard benchmark datasets. Additionally, we demonstrate that our method can generate realistic synthetic images that achieve downstream classification accuracy comparable to SOTA methods. Our method is quite general, as the encoder and decoder modules can be freely substituted to suit different tasks. It is also highly scalable, consisting only of subroutines that scale linearly with the number of samples and/or can be implemented efficiently in distributed systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Private Training & Data Generation by Clustering Embeddings
Zhou, Felix
Zhou, Samson
Mirrokni, Vahab
Epasto, Alessandro
Cohen-Addad, Vincent
Machine Learning
Cryptography and Security
Deep neural networks often use large, high-quality datasets to achieve high performance on many machine learning tasks. When training involves potentially sensitive data, this process can raise privacy concerns, as large models have been shown to unintentionally memorize and reveal sensitive information, including reconstructing entire training samples. Differential privacy (DP) provides a robust framework for protecting individual data and in particular, a new approach to privately training deep neural networks is to approximate the input dataset with a privately generated synthetic dataset, before any subsequent training algorithm. We introduce a novel principled method for DP synthetic image embedding generation, based on fitting a Gaussian Mixture Model (GMM) in an appropriate embedding space using DP clustering. Our method provably learns a GMM under separation conditions. Empirically, a simple two-layer neural network trained on synthetically generated embeddings achieves state-of-the-art (SOTA) classification accuracy on standard benchmark datasets. Additionally, we demonstrate that our method can generate realistic synthetic images that achieve downstream classification accuracy comparable to SOTA methods. Our method is quite general, as the encoder and decoder modules can be freely substituted to suit different tasks. It is also highly scalable, consisting only of subroutines that scale linearly with the number of samples and/or can be implemented efficiently in distributed systems.
title Private Training & Data Generation by Clustering Embeddings
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2506.16661