Secure Retrieval-Augmented Generation against Poisoning Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Zirui, Sun, Jikai, Gao, Anjun, Quan, Yueyang, Liu, Zhuqing, Hu, Xiaohua, Fang, Minghong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912697363202048
author Cheng, Zirui
Sun, Jikai
Gao, Anjun
Quan, Yueyang
Liu, Zhuqing
Hu, Xiaohua
Fang, Minghong
author_facet Cheng, Zirui
Sun, Jikai
Gao, Anjun
Quan, Yueyang
Liu, Zhuqing
Hu, Xiaohua
Fang, Minghong
contents Large language models (LLMs) have transformed natural language processing (NLP), enabling applications from content generation to decision support. Retrieval-Augmented Generation (RAG) improves LLMs by incorporating external knowledge but also introduces security risks, particularly from data poisoning, where the attacker injects poisoned texts into the knowledge database to manipulate system outputs. While various defenses have been proposed, they often struggle against advanced attacks. To address this, we introduce RAGuard, a detection framework designed to identify poisoned texts. RAGuard first expands the retrieval scope to increase the proportion of clean texts, reducing the likelihood of retrieving poisoned content. It then applies chunk-wise perplexity filtering to detect abnormal variations and text similarity filtering to flag highly similar texts. This non-parametric approach enhances RAG security, and experiments on large-scale datasets demonstrate its effectiveness in detecting and mitigating poisoning attacks, including strong adaptive attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Secure Retrieval-Augmented Generation against Poisoning Attacks
Cheng, Zirui
Sun, Jikai
Gao, Anjun
Quan, Yueyang
Liu, Zhuqing
Hu, Xiaohua
Fang, Minghong
Cryptography and Security
Information Retrieval
Machine Learning
Large language models (LLMs) have transformed natural language processing (NLP), enabling applications from content generation to decision support. Retrieval-Augmented Generation (RAG) improves LLMs by incorporating external knowledge but also introduces security risks, particularly from data poisoning, where the attacker injects poisoned texts into the knowledge database to manipulate system outputs. While various defenses have been proposed, they often struggle against advanced attacks. To address this, we introduce RAGuard, a detection framework designed to identify poisoned texts. RAGuard first expands the retrieval scope to increase the proportion of clean texts, reducing the likelihood of retrieving poisoned content. It then applies chunk-wise perplexity filtering to detect abnormal variations and text similarity filtering to flag highly similar texts. This non-parametric approach enhances RAG security, and experiments on large-scale datasets demonstrate its effectiveness in detecting and mitigating poisoning attacks, including strong adaptive attacks.
title Secure Retrieval-Augmented Generation against Poisoning Attacks
topic Cryptography and Security
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2510.25025