SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Xun, Niu, Simin, Li, Zhiyu, Zhang, Sensen, Wang, Hanyu, Xiong, Feiyu, Fan, Jason Zhaoxin, Tang, Bo, Song, Shichao, Wang, Mengwei, Yang, Jiawei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916625327849472
author Liang, Xun
Niu, Simin
Li, Zhiyu
Zhang, Sensen
Wang, Hanyu
Xiong, Feiyu
Fan, Jason Zhaoxin
Tang, Bo
Song, Shichao
Wang, Mengwei
Yang, Jiawei
author_facet Liang, Xun
Niu, Simin
Li, Zhiyu
Zhang, Sensen
Wang, Hanyu
Xiong, Feiyu
Fan, Jason Zhaoxin
Tang, Bo
Song, Shichao
Wang, Mengwei
Yang, Jiawei
contents The indexing-retrieval-generation paradigm of retrieval-augmented generation (RAG) has been highly successful in solving knowledge-intensive tasks by integrating external knowledge into large language models (LLMs). However, the incorporation of external and unverified knowledge increases the vulnerability of LLMs because attackers can perform attack tasks by manipulating knowledge. In this paper, we introduce a benchmark named SafeRAG designed to evaluate the RAG security. First, we classify attack tasks into silver noise, inter-context conflict, soft ad, and white Denial-of-Service. Next, we construct RAG security evaluation dataset (i.e., SafeRAG dataset) primarily manually for each task. We then utilize the SafeRAG dataset to simulate various attack scenarios that RAG may encounter. Experiments conducted on 14 representative RAG components demonstrate that RAG exhibits significant vulnerability to all attack tasks and even the most apparent attack task can easily bypass existing retrievers, filters, or advanced LLMs, resulting in the degradation of RAG service quality. Code is available at: https://github.com/IAAR-Shanghai/SafeRAG.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18636
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model
Liang, Xun
Niu, Simin
Li, Zhiyu
Zhang, Sensen
Wang, Hanyu
Xiong, Feiyu
Fan, Jason Zhaoxin
Tang, Bo
Song, Shichao
Wang, Mengwei
Yang, Jiawei
Cryptography and Security
Artificial Intelligence
Information Retrieval
The indexing-retrieval-generation paradigm of retrieval-augmented generation (RAG) has been highly successful in solving knowledge-intensive tasks by integrating external knowledge into large language models (LLMs). However, the incorporation of external and unverified knowledge increases the vulnerability of LLMs because attackers can perform attack tasks by manipulating knowledge. In this paper, we introduce a benchmark named SafeRAG designed to evaluate the RAG security. First, we classify attack tasks into silver noise, inter-context conflict, soft ad, and white Denial-of-Service. Next, we construct RAG security evaluation dataset (i.e., SafeRAG dataset) primarily manually for each task. We then utilize the SafeRAG dataset to simulate various attack scenarios that RAG may encounter. Experiments conducted on 14 representative RAG components demonstrate that RAG exhibits significant vulnerability to all attack tasks and even the most apparent attack task can easily bypass existing retrievers, filters, or advanced LLMs, resulting in the degradation of RAG service quality. Code is available at: https://github.com/IAAR-Shanghai/SafeRAG.
title SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model
topic Cryptography and Security
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2501.18636