Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Koga, Tatsuki, Wu, Ruihan, Zhang, Zhiyuan, Chaudhuri, Kamalika
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912702668996608
author Koga, Tatsuki
Wu, Ruihan
Zhang, Zhiyuan
Chaudhuri, Kamalika
author_facet Koga, Tatsuki
Wu, Ruihan
Zhang, Zhiyuan
Chaudhuri, Kamalika
contents With the recent remarkable advancement of large language models (LLMs), there has been a growing interest in utilizing them in the domains with highly sensitive data that lies outside their training data. For this purpose, retrieval-augmented generation (RAG) is particularly effective -- it assists LLMs by directly providing relevant information from the external knowledge sources. However, without extra privacy safeguards, RAG outputs risk leaking sensitive information from the external data source. In this work, we explore RAG under differential privacy (DP), a formal guarantee of data privacy. The main challenge with differentially private RAG is how to generate long accurate answers within a moderate privacy budget. We address this by proposing an algorithm that smartly spends privacy budget only for the tokens that require the sensitive information and uses the non-private LLM for other tokens. Our extensive empirical evaluations reveal that our algorithm outperforms the non-RAG baseline under a reasonable privacy budget of $ε\approx 10$ across different models and datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04697
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
Koga, Tatsuki
Wu, Ruihan
Zhang, Zhiyuan
Chaudhuri, Kamalika
Cryptography and Security
Artificial Intelligence
Computation and Language
With the recent remarkable advancement of large language models (LLMs), there has been a growing interest in utilizing them in the domains with highly sensitive data that lies outside their training data. For this purpose, retrieval-augmented generation (RAG) is particularly effective -- it assists LLMs by directly providing relevant information from the external knowledge sources. However, without extra privacy safeguards, RAG outputs risk leaking sensitive information from the external data source. In this work, we explore RAG under differential privacy (DP), a formal guarantee of data privacy. The main challenge with differentially private RAG is how to generate long accurate answers within a moderate privacy budget. We address this by proposing an algorithm that smartly spends privacy budget only for the tokens that require the sensitive information and uses the non-private LLM for other tokens. Our extensive empirical evaluations reveal that our algorithm outperforms the non-RAG baseline under a reasonable privacy budget of $ε\approx 10$ across different models and datasets.
title Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.04697