SynBench: A Benchmark for Differentially Private Text Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Yidan, Schlegel, Viktor, Nandakumar, Srinivasan, Zahid, Iqra, Wu, Yuping, Wu, Yulong, Li, Hao, Zhang, Jie, Del-Pinto, Warren, Nenadic, Goran, Lam, Siew Kei, Bharath, Anil Anthony
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915985128161280
author Sun, Yidan
Schlegel, Viktor
Nandakumar, Srinivasan
Zahid, Iqra
Wu, Yuping
Wu, Yulong
Li, Hao
Zhang, Jie
Del-Pinto, Warren
Nenadic, Goran
Lam, Siew Kei
Bharath, Anil Anthony
author_facet Sun, Yidan
Schlegel, Viktor
Nandakumar, Srinivasan
Zahid, Iqra
Wu, Yuping
Wu, Yulong
Li, Hao
Zhang, Jie
Del-Pinto, Warren
Nenadic, Goran
Lam, Siew Kei
Bharath, Anil Anthony
contents Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regulatory boundaries, while bounding the risks of re-identification and membership inference. LLM-based methods deliver promising results; however, comparisons are exacerbated by differing evaluation setups and "private" datasets, potential pre-training contamination is not considered and guarantees are not verified with DP audits. To advance this field, we introduce a unified evaluation framework with standardised utility and fidelity metrics and privacy audits, encompassing nine curated datasets that capture domain-specific complexities such as technical jargon, long-context dependencies, and specialised document structures. In a large-scale empirical study, we benchmark LLM-based state-of-the-art DP text generators of varying sizes (between 1--8B). Our results indicate that DP synthetic text generation remains an unsolved challenge, with quality deteriorating more as the private datasets deviate further from the generators' pre-training corpora. Our novel synthetic text membership inference attack (MIA) explains this observation: Synthetic data quality is overestimated when LLMs have been pre-trained -- without DP -- on portions of the "private" data to be generated. Finally, our work provides the first quantitative evidence that this "public pre-training and private generation" paradigm invalidates the guaranteed privacy bounds of real-world private datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SynBench: A Benchmark for Differentially Private Text Generation
Sun, Yidan
Schlegel, Viktor
Nandakumar, Srinivasan
Zahid, Iqra
Wu, Yuping
Wu, Yulong
Li, Hao
Zhang, Jie
Del-Pinto, Warren
Nenadic, Goran
Lam, Siew Kei
Bharath, Anil Anthony
Artificial Intelligence
Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regulatory boundaries, while bounding the risks of re-identification and membership inference. LLM-based methods deliver promising results; however, comparisons are exacerbated by differing evaluation setups and "private" datasets, potential pre-training contamination is not considered and guarantees are not verified with DP audits. To advance this field, we introduce a unified evaluation framework with standardised utility and fidelity metrics and privacy audits, encompassing nine curated datasets that capture domain-specific complexities such as technical jargon, long-context dependencies, and specialised document structures. In a large-scale empirical study, we benchmark LLM-based state-of-the-art DP text generators of varying sizes (between 1--8B). Our results indicate that DP synthetic text generation remains an unsolved challenge, with quality deteriorating more as the private datasets deviate further from the generators' pre-training corpora. Our novel synthetic text membership inference attack (MIA) explains this observation: Synthetic data quality is overestimated when LLMs have been pre-trained -- without DP -- on portions of the "private" data to be generated. Finally, our work provides the first quantitative evidence that this "public pre-training and private generation" paradigm invalidates the guaranteed privacy bounds of real-world private datasets.
title SynBench: A Benchmark for Differentially Private Text Generation
topic Artificial Intelligence
url https://arxiv.org/abs/2509.14594