Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Atwal, Tevin, Tieu, Chan Nam, Yuan, Yefeng, Shi, Zhan, Liu, Yuhong, Cheng, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916861265838080
author Atwal, Tevin
Tieu, Chan Nam
Yuan, Yefeng
Shi, Zhan
Liu, Yuhong
Cheng, Liang
author_facet Atwal, Tevin
Tieu, Chan Nam
Yuan, Yefeng
Shi, Zhan
Liu, Yuhong
Cheng, Liang
contents The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data to facilitate model training, its diversity and privacy risks remain underexplored. Focusing on text-based synthetic data, we propose a comprehensive set of metrics to quantitatively assess the diversity (i.e., linguistic expression, sentiment, and user perspective), and privacy (i.e., re-identification risk and stylistic outliers) of synthetic datasets generated by several state-of-the-art LLMs. Experiment results reveal significant limitations in LLMs' capabilities in generating diverse and privacy-preserving synthetic data. Guided by the evaluation results, a prompt-based approach is proposed to enhance the diversity of synthetic reviews while preserving reviewer privacy.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
Atwal, Tevin
Tieu, Chan Nam
Yuan, Yefeng
Shi, Zhan
Liu, Yuhong
Cheng, Liang
Computation and Language
Cryptography and Security
Machine Learning
The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data to facilitate model training, its diversity and privacy risks remain underexplored. Focusing on text-based synthetic data, we propose a comprehensive set of metrics to quantitatively assess the diversity (i.e., linguistic expression, sentiment, and user perspective), and privacy (i.e., re-identification risk and stylistic outliers) of synthetic datasets generated by several state-of-the-art LLMs. Experiment results reveal significant limitations in LLMs' capabilities in generating diverse and privacy-preserving synthetic data. Guided by the evaluation results, a prompt-based approach is proposed to enhance the diversity of synthetic reviews while preserving reviewer privacy.
title Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
topic Computation and Language
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2507.18055