A Consensus Privacy Metrics Framework for Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pilgram, Lisa, Dankar, Fida K., Drechsler, Jorg, Elliot, Mark, Domingo-Ferrer, Josep, Francis, Paul, Kantarcioglu, Murat, Kong, Linglong, Malin, Bradley, Muralidhar, Krishnamurty, Myles, Puja, Prasser, Fabian, Raisaro, Jean Louis, Yan, Chao, Emam, Khaled El
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911622067388416
author Pilgram, Lisa
Dankar, Fida K.
Drechsler, Jorg
Elliot, Mark
Domingo-Ferrer, Josep
Francis, Paul
Kantarcioglu, Murat
Kong, Linglong
Malin, Bradley
Muralidhar, Krishnamurty
Myles, Puja
Prasser, Fabian
Raisaro, Jean Louis
Yan, Chao
Emam, Khaled El
author_facet Pilgram, Lisa
Dankar, Fida K.
Drechsler, Jorg
Elliot, Mark
Domingo-Ferrer, Josep
Francis, Paul
Kantarcioglu, Murat
Kong, Linglong
Malin, Bradley
Muralidhar, Krishnamurty
Myles, Puja
Prasser, Fabian
Raisaro, Jean Louis
Yan, Chao
Emam, Khaled El
contents Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04980
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Consensus Privacy Metrics Framework for Synthetic Data
Pilgram, Lisa
Dankar, Fida K.
Drechsler, Jorg
Elliot, Mark
Domingo-Ferrer, Josep
Francis, Paul
Kantarcioglu, Murat
Kong, Linglong
Malin, Bradley
Muralidhar, Krishnamurty
Myles, Puja
Prasser, Fabian
Raisaro, Jean Louis
Yan, Chao
Emam, Khaled El
Cryptography and Security
Artificial Intelligence
Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.
title A Consensus Privacy Metrics Framework for Synthetic Data
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2503.04980