Benchmarking Poisoning Attacks against Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Baolei, Xin, Haoran, Li, Jiatong, Zhang, Dongzhe, Fang, Minghong, Liu, Zhuqing, Nie, Lihai, Liu, Zheli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913857546485760
author Zhang, Baolei
Xin, Haoran
Li, Jiatong
Zhang, Dongzhe
Fang, Minghong
Liu, Zhuqing
Nie, Lihai
Liu, Zheli
author_facet Zhang, Baolei
Xin, Haoran
Li, Jiatong
Zhang, Dongzhe
Fang, Minghong
Liu, Zhuqing
Nie, Lihai
Liu, Zheli
contents Retrieval-Augmented Generation (RAG) has proven effective in mitigating hallucinations in large language models by incorporating external knowledge during inference. However, this integration introduces new security vulnerabilities, particularly to poisoning attacks. Although prior work has explored various poisoning strategies, a thorough assessment of their practical threat to RAG systems remains missing. To address this gap, we propose the first comprehensive benchmark framework for evaluating poisoning attacks on RAG. Our benchmark covers 5 standard question answering (QA) datasets and 10 expanded variants, along with 13 poisoning attack methods and 7 defense mechanisms, representing a broad spectrum of existing techniques. Using this benchmark, we conduct a comprehensive evaluation of all included attacks and defenses across the full dataset spectrum. Our findings show that while existing attacks perform well on standard QA datasets, their effectiveness drops significantly on the expanded versions. Moreover, our results demonstrate that various advanced RAG architectures, such as sequential, branching, conditional, and loop RAG, as well as multi-turn conversational RAG, multimodal RAG systems, and RAG-based LLM agent systems, remain susceptible to poisoning attacks. Notably, current defense techniques fail to provide robust protection, underscoring the pressing need for more resilient and generalizable defense strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Poisoning Attacks against Retrieval-Augmented Generation
Zhang, Baolei
Xin, Haoran
Li, Jiatong
Zhang, Dongzhe
Fang, Minghong
Liu, Zhuqing
Nie, Lihai
Liu, Zheli
Cryptography and Security
Information Retrieval
Machine Learning
Retrieval-Augmented Generation (RAG) has proven effective in mitigating hallucinations in large language models by incorporating external knowledge during inference. However, this integration introduces new security vulnerabilities, particularly to poisoning attacks. Although prior work has explored various poisoning strategies, a thorough assessment of their practical threat to RAG systems remains missing. To address this gap, we propose the first comprehensive benchmark framework for evaluating poisoning attacks on RAG. Our benchmark covers 5 standard question answering (QA) datasets and 10 expanded variants, along with 13 poisoning attack methods and 7 defense mechanisms, representing a broad spectrum of existing techniques. Using this benchmark, we conduct a comprehensive evaluation of all included attacks and defenses across the full dataset spectrum. Our findings show that while existing attacks perform well on standard QA datasets, their effectiveness drops significantly on the expanded versions. Moreover, our results demonstrate that various advanced RAG architectures, such as sequential, branching, conditional, and loop RAG, as well as multi-turn conversational RAG, multimodal RAG systems, and RAG-based LLM agent systems, remain susceptible to poisoning attacks. Notably, current defense techniques fail to provide robust protection, underscoring the pressing need for more resilient and generalizable defense strategies.
title Benchmarking Poisoning Attacks against Retrieval-Augmented Generation
topic Cryptography and Security
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2505.18543