End to End Collaborative Synthetic Data Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pentyala, Sikha, Sitaraman, Geetha, Claar, Trae, De Cock, Martine
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917267556532224
author Pentyala, Sikha
Sitaraman, Geetha
Claar, Trae
De Cock, Martine
author_facet Pentyala, Sikha
Sitaraman, Geetha
Claar, Trae
De Cock, Martine
contents The success of AI is based on the availability of data to train models. While in some cases a single data custodian may have sufficient data to enable AI, often multiple custodians need to collaborate to reach a cumulative size required for meaningful AI research. The latter is, for example, often the case for rare diseases, with each clinical site having data for only a small number of patients. Recent algorithms for federated synthetic data generation are an important step towards collaborative, privacy-preserving data sharing. Existing techniques, however, focus exclusively on synthesizer training, assuming that the training data is already preprocessed and that the desired synthetic data can be delivered in one shot, without any hyperparameter tuning. In this paper, we propose an end-to-end collaborative framework for publishing of synthetic data that accounts for privacy-preserving preprocessing as well as evaluation. We instantiate this framework with Secure Multiparty Computation (MPC) protocols and evaluate it in a use case for privacy-preserving publishing of synthetic genomic data for leukemia.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03766
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle End to End Collaborative Synthetic Data Generation
Pentyala, Sikha
Sitaraman, Geetha
Claar, Trae
De Cock, Martine
Cryptography and Security
Machine Learning
The success of AI is based on the availability of data to train models. While in some cases a single data custodian may have sufficient data to enable AI, often multiple custodians need to collaborate to reach a cumulative size required for meaningful AI research. The latter is, for example, often the case for rare diseases, with each clinical site having data for only a small number of patients. Recent algorithms for federated synthetic data generation are an important step towards collaborative, privacy-preserving data sharing. Existing techniques, however, focus exclusively on synthesizer training, assuming that the training data is already preprocessed and that the desired synthetic data can be delivered in one shot, without any hyperparameter tuning. In this paper, we propose an end-to-end collaborative framework for publishing of synthetic data that accounts for privacy-preserving preprocessing as well as evaluation. We instantiate this framework with Secure Multiparty Computation (MPC) protocols and evaluate it in a use case for privacy-preserving publishing of synthetic genomic data for leukemia.
title End to End Collaborative Synthetic Data Generation
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2412.03766