Privacy-preserving datasets by capturing feature distributions with Conditional VAEs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Di Salvo, Francesco, Tafler, David, Doerrich, Sebastian, Ledig, Christian
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913455378792448
author Di Salvo, Francesco
Tafler, David
Doerrich, Sebastian
Ledig, Christian
author_facet Di Salvo, Francesco
Tafler, David
Doerrich, Sebastian
Ledig, Christian
contents Large and well-annotated datasets are essential for advancing deep learning applications, however often costly or impossible to obtain by a single entity. In many areas, including the medical domain, approaches relying on data sharing have become critical to address those challenges. While effective in increasing dataset size and diversity, data sharing raises significant privacy concerns. Commonly employed anonymization methods based on the k-anonymity paradigm often fail to preserve data diversity, affecting model robustness. This work introduces a novel approach using Conditional Variational Autoencoders (CVAEs) trained on feature vectors extracted from large pre-trained vision foundation models. Foundation models effectively detect and represent complex patterns across diverse domains, allowing the CVAE to faithfully capture the embedding space of a given data distribution to generate (sample) a diverse, privacy-respecting, and potentially unbounded set of synthetic feature vectors. Our method notably outperforms traditional approaches in both medical and natural image domains, exhibiting greater dataset diversity and higher robustness against perturbations while preserving sample privacy. These results underscore the potential of generative models to significantly impact deep learning applications in data-scarce and privacy-sensitive environments. The source code is available at https://github.com/francescodisalvo05/cvae-anonymization .
format Preprint
id arxiv_https___arxiv_org_abs_2408_00639
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Privacy-preserving datasets by capturing feature distributions with Conditional VAEs
Di Salvo, Francesco
Tafler, David
Doerrich, Sebastian
Ledig, Christian
Machine Learning
Computer Vision and Pattern Recognition
Image and Video Processing
Large and well-annotated datasets are essential for advancing deep learning applications, however often costly or impossible to obtain by a single entity. In many areas, including the medical domain, approaches relying on data sharing have become critical to address those challenges. While effective in increasing dataset size and diversity, data sharing raises significant privacy concerns. Commonly employed anonymization methods based on the k-anonymity paradigm often fail to preserve data diversity, affecting model robustness. This work introduces a novel approach using Conditional Variational Autoencoders (CVAEs) trained on feature vectors extracted from large pre-trained vision foundation models. Foundation models effectively detect and represent complex patterns across diverse domains, allowing the CVAE to faithfully capture the embedding space of a given data distribution to generate (sample) a diverse, privacy-respecting, and potentially unbounded set of synthetic feature vectors. Our method notably outperforms traditional approaches in both medical and natural image domains, exhibiting greater dataset diversity and higher robustness against perturbations while preserving sample privacy. These results underscore the potential of generative models to significantly impact deep learning applications in data-scarce and privacy-sensitive environments. The source code is available at https://github.com/francescodisalvo05/cvae-anonymization .
title Privacy-preserving datasets by capturing feature distributions with Conditional VAEs
topic Machine Learning
Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2408.00639