Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xie, Yuanbo, Zhang, Yingjie, Li, Yulin, Song, Shouyou, Chen, Xiaokun, Liu, Zhihan, Su, Liya, Liu, Tingwen
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913025158545408
author Xie, Yuanbo
Zhang, Yingjie
Li, Yulin
Song, Shouyou
Chen, Xiaokun
Liu, Zhihan
Su, Liya
Liu, Tingwen
author_facet Xie, Yuanbo
Zhang, Yingjie
Li, Yulin
Song, Shouyou
Chen, Xiaokun
Liu, Zhihan
Su, Liya
Liu, Tingwen
contents Retrieval-Augmented Generation (RAG) systems augment large language models with external knowledge, yet introduce a critical security vulnerability: RAG Knowledge Base Leakage, wherein adversarial prompts can induce the model to divulge retrieved proprietary content. Recent studies reveal that such leakage can be executed through adaptive and iterative attack strategies (named RAG extraction attack), while effective countermeasures remain notably lacking. To bridge this gap, we propose CanaryRAG, a runtime defense mechanism inspired by stack canaries in software security. CanaryRAG embeds carefully designed canary tokens into retrieved chunks and reformulates RAG extraction defense as a dual-path runtime integrity game. Leakage is detected in real time whenever either the target or oracle path violates its expected canary behavior, including under adaptive suppression and obfuscation. Extensive evaluations against existing attacks demonstrate that CanaryRAG provides robust defense, achieving substantially lower chunk recovery rates than state-of-the-art baselines while imposing negligible impact on task performance and inference latency. Moreover, as a plug-and-play solution, CanaryRAG can be seamlessly integrated into arbitrary RAG pipelines without requiring retraining or structural modifications, offering a practical and scalable safeguard for proprietary data.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10717
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game
Xie, Yuanbo
Zhang, Yingjie
Li, Yulin
Song, Shouyou
Chen, Xiaokun
Liu, Zhihan
Su, Liya
Liu, Tingwen
Cryptography and Security
Artificial Intelligence
Computation and Language
Retrieval-Augmented Generation (RAG) systems augment large language models with external knowledge, yet introduce a critical security vulnerability: RAG Knowledge Base Leakage, wherein adversarial prompts can induce the model to divulge retrieved proprietary content. Recent studies reveal that such leakage can be executed through adaptive and iterative attack strategies (named RAG extraction attack), while effective countermeasures remain notably lacking. To bridge this gap, we propose CanaryRAG, a runtime defense mechanism inspired by stack canaries in software security. CanaryRAG embeds carefully designed canary tokens into retrieved chunks and reformulates RAG extraction defense as a dual-path runtime integrity game. Leakage is detected in real time whenever either the target or oracle path violates its expected canary behavior, including under adaptive suppression and obfuscation. Extensive evaluations against existing attacks demonstrate that CanaryRAG provides robust defense, achieving substantially lower chunk recovery rates than state-of-the-art baselines while imposing negligible impact on task performance and inference latency. Moreover, as a plug-and-play solution, CanaryRAG can be seamlessly integrated into arbitrary RAG pipelines without requiring retraining or structural modifications, offering a practical and scalable safeguard for proprietary data.
title Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.10717