SynBench: A Benchmark for Differentially Private Text Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915985128161280 |
|---|---|
| author | Sun, Yidan Schlegel, Viktor Nandakumar, Srinivasan Zahid, Iqra Wu, Yuping Wu, Yulong Li, Hao Zhang, Jie Del-Pinto, Warren Nenadic, Goran Lam, Siew Kei Bharath, Anil Anthony |
| author_facet | Sun, Yidan Schlegel, Viktor Nandakumar, Srinivasan Zahid, Iqra Wu, Yuping Wu, Yulong Li, Hao Zhang, Jie Del-Pinto, Warren Nenadic, Goran Lam, Siew Kei Bharath, Anil Anthony |
| contents | Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regulatory boundaries, while bounding the risks of re-identification and membership inference. LLM-based methods deliver promising results; however, comparisons are exacerbated by differing evaluation setups and "private" datasets, potential pre-training contamination is not considered and guarantees are not verified with DP audits. To advance this field, we introduce a unified evaluation framework with standardised utility and fidelity metrics and privacy audits, encompassing nine curated datasets that capture domain-specific complexities such as technical jargon, long-context dependencies, and specialised document structures. In a large-scale empirical study, we benchmark LLM-based state-of-the-art DP text generators of varying sizes (between 1--8B). Our results indicate that DP synthetic text generation remains an unsolved challenge, with quality deteriorating more as the private datasets deviate further from the generators' pre-training corpora. Our novel synthetic text membership inference attack (MIA) explains this observation: Synthetic data quality is overestimated when LLMs have been pre-trained -- without DP -- on portions of the "private" data to be generated. Finally, our work provides the first quantitative evidence that this "public pre-training and private generation" paradigm invalidates the guaranteed privacy bounds of real-world private datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14594 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SynBench: A Benchmark for Differentially Private Text Generation Sun, Yidan Schlegel, Viktor Nandakumar, Srinivasan Zahid, Iqra Wu, Yuping Wu, Yulong Li, Hao Zhang, Jie Del-Pinto, Warren Nenadic, Goran Lam, Siew Kei Bharath, Anil Anthony Artificial Intelligence Synthetic text generation with Differential Privacy (DP) guarantees emerges as a principled approach that can enable the sharing of sensitive datasets across institutional and regulatory boundaries, while bounding the risks of re-identification and membership inference. LLM-based methods deliver promising results; however, comparisons are exacerbated by differing evaluation setups and "private" datasets, potential pre-training contamination is not considered and guarantees are not verified with DP audits. To advance this field, we introduce a unified evaluation framework with standardised utility and fidelity metrics and privacy audits, encompassing nine curated datasets that capture domain-specific complexities such as technical jargon, long-context dependencies, and specialised document structures. In a large-scale empirical study, we benchmark LLM-based state-of-the-art DP text generators of varying sizes (between 1--8B). Our results indicate that DP synthetic text generation remains an unsolved challenge, with quality deteriorating more as the private datasets deviate further from the generators' pre-training corpora. Our novel synthetic text membership inference attack (MIA) explains this observation: Synthetic data quality is overestimated when LLMs have been pre-trained -- without DP -- on portions of the "private" data to be generated. Finally, our work provides the first quantitative evidence that this "public pre-training and private generation" paradigm invalidates the guaranteed privacy bounds of real-world private datasets. |
| title | SynBench: A Benchmark for Differentially Private Text Generation |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2509.14594 |