Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916861265838080 |
|---|---|
| author | Atwal, Tevin Tieu, Chan Nam Yuan, Yefeng Shi, Zhan Liu, Yuhong Cheng, Liang |
| author_facet | Atwal, Tevin Tieu, Chan Nam Yuan, Yefeng Shi, Zhan Liu, Yuhong Cheng, Liang |
| contents | The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data to facilitate model training, its diversity and privacy risks remain underexplored. Focusing on text-based synthetic data, we propose a comprehensive set of metrics to quantitatively assess the diversity (i.e., linguistic expression, sentiment, and user perspective), and privacy (i.e., re-identification risk and stylistic outliers) of synthetic datasets generated by several state-of-the-art LLMs. Experiment results reveal significant limitations in LLMs' capabilities in generating diverse and privacy-preserving synthetic data. Guided by the evaluation results, a prompt-based approach is proposed to enhance the diversity of synthetic reviews while preserving reviewer privacy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_18055 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs Atwal, Tevin Tieu, Chan Nam Yuan, Yefeng Shi, Zhan Liu, Yuhong Cheng, Liang Computation and Language Cryptography and Security Machine Learning The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data to facilitate model training, its diversity and privacy risks remain underexplored. Focusing on text-based synthetic data, we propose a comprehensive set of metrics to quantitatively assess the diversity (i.e., linguistic expression, sentiment, and user perspective), and privacy (i.e., re-identification risk and stylistic outliers) of synthetic datasets generated by several state-of-the-art LLMs. Experiment results reveal significant limitations in LLMs' capabilities in generating diverse and privacy-preserving synthetic data. Guided by the evaluation results, a prompt-based approach is proposed to enhance the diversity of synthetic reviews while preserving reviewer privacy. |
| title | Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs |
| topic | Computation and Language Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2507.18055 |