SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shanmugarasa, Yashothara, Ding, Ming, Chamikara, M. A. P, Rakotoarivelo, Thierry
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909652860534784
author Shanmugarasa, Yashothara
Ding, Ming
Chamikara, M. A. P
Rakotoarivelo, Thierry
author_facet Shanmugarasa, Yashothara
Ding, Ming
Chamikara, M. A. P
Rakotoarivelo, Thierry
contents Large language models (LLMs) are sophisticated artificial intelligence systems that enable machines to generate human-like text with remarkable precision. While LLMs offer significant technological progress, their development using vast amounts of user data scraped from the web and collected from extensive user interactions poses risks of sensitive information leakage. Most existing surveys focus on the privacy implications of the training data but tend to overlook privacy risks from user interactions and advanced LLM capabilities. This paper aims to fill that gap by providing a comprehensive analysis of privacy in LLMs, categorizing the challenges into four main areas: (i) privacy issues in LLM training data, (ii) privacy challenges associated with user prompts, (iii) privacy vulnerabilities in LLM-generated outputs, and (iv) privacy challenges involving LLM agents. We evaluate the effectiveness and limitations of existing mitigation mechanisms targeting these proposed privacy challenges and identify areas for further research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
Shanmugarasa, Yashothara
Ding, Ming
Chamikara, M. A. P
Rakotoarivelo, Thierry
Cryptography and Security
Human-Computer Interaction
Large language models (LLMs) are sophisticated artificial intelligence systems that enable machines to generate human-like text with remarkable precision. While LLMs offer significant technological progress, their development using vast amounts of user data scraped from the web and collected from extensive user interactions poses risks of sensitive information leakage. Most existing surveys focus on the privacy implications of the training data but tend to overlook privacy risks from user interactions and advanced LLM capabilities. This paper aims to fill that gap by providing a comprehensive analysis of privacy in LLMs, categorizing the challenges into four main areas: (i) privacy issues in LLM training data, (ii) privacy challenges associated with user prompts, (iii) privacy vulnerabilities in LLM-generated outputs, and (iv) privacy challenges involving LLM agents. We evaluate the effectiveness and limitations of existing mitigation mechanisms targeting these proposed privacy challenges and identify areas for further research.
title SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
topic Cryptography and Security
Human-Computer Interaction
url https://arxiv.org/abs/2506.12699