On Protecting the Data Privacy of Large Language Models (LLMs): A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Biwei, Li, Kun, Xu, Minghui, Dong, Yueyan, Zhang, Yue, Ren, Zhaochun, Cheng, Xiuzhen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929275741929472
author Yan, Biwei
Li, Kun
Xu, Minghui
Dong, Yueyan
Zhang, Yue
Ren, Zhaochun
Cheng, Xiuzhen
author_facet Yan, Biwei
Li, Kun
Xu, Minghui
Dong, Yueyan
Zhang, Yue
Ren, Zhaochun
Cheng, Xiuzhen
contents Large language models (LLMs) are complex artificial intelligence systems capable of understanding, generating and translating human language. They learn language patterns by analyzing large amounts of text data, allowing them to perform writing, conversation, summarizing and other language tasks. When LLMs process and generate large amounts of data, there is a risk of leaking sensitive information, which may threaten data privacy. This paper concentrates on elucidating the data privacy concerns associated with LLMs to foster a comprehensive understanding. Specifically, a thorough investigation is undertaken to delineate the spectrum of data privacy threats, encompassing both passive privacy leakage and active privacy attacks within LLMs. Subsequently, we conduct an assessment of the privacy protection mechanisms employed by LLMs at various stages, followed by a detailed examination of their efficacy and constraints. Finally, the discourse extends to delineate the challenges encountered and outline prospective directions for advancement in the realm of LLM privacy protection.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05156
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
Yan, Biwei
Li, Kun
Xu, Minghui
Dong, Yueyan
Zhang, Yue
Ren, Zhaochun
Cheng, Xiuzhen
Cryptography and Security
Large language models (LLMs) are complex artificial intelligence systems capable of understanding, generating and translating human language. They learn language patterns by analyzing large amounts of text data, allowing them to perform writing, conversation, summarizing and other language tasks. When LLMs process and generate large amounts of data, there is a risk of leaking sensitive information, which may threaten data privacy. This paper concentrates on elucidating the data privacy concerns associated with LLMs to foster a comprehensive understanding. Specifically, a thorough investigation is undertaken to delineate the spectrum of data privacy threats, encompassing both passive privacy leakage and active privacy attacks within LLMs. Subsequently, we conduct an assessment of the privacy protection mechanisms employed by LLMs at various stages, followed by a detailed examination of their efficacy and constraints. Finally, the discourse extends to delineate the challenges encountered and outline prospective directions for advancement in the realm of LLM privacy protection.
title On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
topic Cryptography and Security
url https://arxiv.org/abs/2403.05156