The Silent Saboteur: Imperceptible Adversarial Attacks against Black-Box Retrieval-Augmented Generation Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Hongru, Liu, Yu-an, Zhang, Ruqing, Guo, Jiafeng, Lv, Jianming, de Rijke, Maarten, Cheng, Xueqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910972433661952
author Song, Hongru
Liu, Yu-an
Zhang, Ruqing
Guo, Jiafeng
Lv, Jianming
de Rijke, Maarten
Cheng, Xueqi
author_facet Song, Hongru
Liu, Yu-an
Zhang, Ruqing
Guo, Jiafeng
Lv, Jianming
de Rijke, Maarten
Cheng, Xueqi
contents We explore adversarial attacks against retrieval-augmented generation (RAG) systems to identify their vulnerabilities. We focus on generating human-imperceptible adversarial examples and introduce a novel imperceptible retrieve-to-generate attack against RAG. This task aims to find imperceptible perturbations that retrieve a target document, originally excluded from the initial top-$k$ candidate set, in order to influence the final answer generation. To address this task, we propose ReGENT, a reinforcement learning-based framework that tracks interactions between the attacker and the target RAG and continuously refines attack strategies based on relevance-generation-naturalness rewards. Experiments on newly constructed factual and non-factual question-answering benchmarks demonstrate that ReGENT significantly outperforms existing attack methods in misleading RAG systems with small imperceptible text perturbations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18583
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Silent Saboteur: Imperceptible Adversarial Attacks against Black-Box Retrieval-Augmented Generation Systems
Song, Hongru
Liu, Yu-an
Zhang, Ruqing
Guo, Jiafeng
Lv, Jianming
de Rijke, Maarten
Cheng, Xueqi
Information Retrieval
We explore adversarial attacks against retrieval-augmented generation (RAG) systems to identify their vulnerabilities. We focus on generating human-imperceptible adversarial examples and introduce a novel imperceptible retrieve-to-generate attack against RAG. This task aims to find imperceptible perturbations that retrieve a target document, originally excluded from the initial top-$k$ candidate set, in order to influence the final answer generation. To address this task, we propose ReGENT, a reinforcement learning-based framework that tracks interactions between the attacker and the target RAG and continuously refines attack strategies based on relevance-generation-naturalness rewards. Experiments on newly constructed factual and non-factual question-answering benchmarks demonstrate that ReGENT significantly outperforms existing attack methods in misleading RAG systems with small imperceptible text perturbations.
title The Silent Saboteur: Imperceptible Adversarial Attacks against Black-Box Retrieval-Augmented Generation Systems
topic Information Retrieval
url https://arxiv.org/abs/2505.18583