A Consensus Privacy Metrics Framework for Synthetic Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911622067388416 |
|---|---|
| author | Pilgram, Lisa Dankar, Fida K. Drechsler, Jorg Elliot, Mark Domingo-Ferrer, Josep Francis, Paul Kantarcioglu, Murat Kong, Linglong Malin, Bradley Muralidhar, Krishnamurty Myles, Puja Prasser, Fabian Raisaro, Jean Louis Yan, Chao Emam, Khaled El |
| author_facet | Pilgram, Lisa Dankar, Fida K. Drechsler, Jorg Elliot, Mark Domingo-Ferrer, Josep Francis, Paul Kantarcioglu, Murat Kong, Linglong Malin, Bradley Muralidhar, Krishnamurty Myles, Puja Prasser, Fabian Raisaro, Jean Louis Yan, Chao Emam, Khaled El |
| contents | Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_04980 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Consensus Privacy Metrics Framework for Synthetic Data Pilgram, Lisa Dankar, Fida K. Drechsler, Jorg Elliot, Mark Domingo-Ferrer, Josep Francis, Paul Kantarcioglu, Murat Kong, Linglong Malin, Bradley Muralidhar, Krishnamurty Myles, Puja Prasser, Fabian Raisaro, Jean Louis Yan, Chao Emam, Khaled El Cryptography and Security Artificial Intelligence Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data. |
| title | A Consensus Privacy Metrics Framework for Synthetic Data |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2503.04980 |