SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Bing, Li, Yixin, Bahamyirou, Asma, Chen, Helen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914261229371392
author Hu, Bing
Li, Yixin
Bahamyirou, Asma
Chen, Helen
author_facet Hu, Bing
Li, Yixin
Bahamyirou, Asma
Chen, Helen
contents The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets for evaluating privacy risks, due to difficulties in acquiring sensitive data. To address this, we introduce SynQP, an open framework for benchmarking privacy in synthetic data generation (SDG) using simulated sensitive data, ensuring that original data remains confidential. We also highlight the need for privacy metrics that fairly account for the probabilistic nature of machine learning models. As a demonstration, we use SynQP to benchmark CTGAN and propose a new identity disclosure risk metric that offers a more accurate estimation of privacy risks compared to existing approaches. Our work provides a critical tool for improving the transparency and reliability of privacy evaluations, enabling safer use of synthetic data in health-related applications. % In our quality evaluations, non-private models achieved near-perfect machine-learning efficacy \(\ge0.97\). Our privacy assessments (Table II) reveal that DP consistently lowers both identity disclosure risk (SD-IDR) and membership-inference attack risk (SD-MIA), with all DP-augmented models staying below the 0.09 regulatory threshold. Code available at https://github.com/CAN-SYNH/SynQP
format Preprint
id arxiv_https___arxiv_org_abs_2601_12124
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
Hu, Bing
Li, Yixin
Bahamyirou, Asma
Chen, Helen
Machine Learning
Artificial Intelligence
The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets for evaluating privacy risks, due to difficulties in acquiring sensitive data. To address this, we introduce SynQP, an open framework for benchmarking privacy in synthetic data generation (SDG) using simulated sensitive data, ensuring that original data remains confidential. We also highlight the need for privacy metrics that fairly account for the probabilistic nature of machine learning models. As a demonstration, we use SynQP to benchmark CTGAN and propose a new identity disclosure risk metric that offers a more accurate estimation of privacy risks compared to existing approaches. Our work provides a critical tool for improving the transparency and reliability of privacy evaluations, enabling safer use of synthetic data in health-related applications. % In our quality evaluations, non-private models achieved near-perfect machine-learning efficacy \(\ge0.97\). Our privacy assessments (Table II) reveal that DP consistently lowers both identity disclosure risk (SD-IDR) and membership-inference attack risk (SD-MIA), with all DP-augmented models staying below the 0.09 regulatory threshold. Code available at https://github.com/CAN-SYNH/SynQP
title SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.12124