LeakDojo: Decoding the Leakage Threats of RAG Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Maosen, Dong, Jianshuo, Lu, Boting, Li, Wenyue, Zhang, Xiaoping, Zhang, Tianwei, Qiu, Han
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909020452814848
author Zhang, Maosen
Dong, Jianshuo
Lu, Boting
Li, Wenyue
Zhang, Xiaoping
Zhang, Tianwei
Qiu, Han
author_facet Zhang, Maosen
Dong, Jianshuo
Lu, Boting
Li, Wenyue
Zhang, Xiaoping
Zhang, Tianwei
Qiu, Han
contents Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to leverage external knowledge, but also exposes valuable RAG databases to leakage attacks. As RAG systems grow more complex and LLMs exhibit stronger instruction-following capabilities, existing studies fall short of systematically assessing RAG leakage risks. We present LeakDojo, a configurable framework for controlled evaluation of RAG leakage. Using LeakDojo, we benchmark six existing attacks across fourteen LLMs, four datasets, and diverse RAG systems. Our study reveals that (1) query generation and adversarial instructions contribute independently to leakage, with overall leakage well approximated by their product; (2) stronger instruction-following capability correlates with higher leakage risk; and (3) improvements in RAG faithfulness can introduce increased leakage risk. These findings provide actionable insights for understanding and mitigating RAG leakage in practice. Our codebase is available at https://github.com/yeasen-z/LeakDojo.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05818
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LeakDojo: Decoding the Leakage Threats of RAG Systems
Zhang, Maosen
Dong, Jianshuo
Lu, Boting
Li, Wenyue
Zhang, Xiaoping
Zhang, Tianwei
Qiu, Han
Cryptography and Security
Artificial Intelligence
Computation and Language
Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to leverage external knowledge, but also exposes valuable RAG databases to leakage attacks. As RAG systems grow more complex and LLMs exhibit stronger instruction-following capabilities, existing studies fall short of systematically assessing RAG leakage risks. We present LeakDojo, a configurable framework for controlled evaluation of RAG leakage. Using LeakDojo, we benchmark six existing attacks across fourteen LLMs, four datasets, and diverse RAG systems. Our study reveals that (1) query generation and adversarial instructions contribute independently to leakage, with overall leakage well approximated by their product; (2) stronger instruction-following capability correlates with higher leakage risk; and (3) improvements in RAG faithfulness can introduce increased leakage risk. These findings provide actionable insights for understanding and mitigating RAG leakage in practice. Our codebase is available at https://github.com/yeasen-z/LeakDojo.
title LeakDojo: Decoding the Leakage Threats of RAG Systems
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.05818