MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Tailun, He, Yu, Wang, Yan, Shao, Shuo, Zheng, Haolun, Liu, Zhihao, Li, Jinfeng, Qin, Zhizhen, Chen, Yuefeng, Chu, Zhixuan, Qin, Zhan, Ren, Kui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914262599860224
author Chen, Tailun
He, Yu
Wang, Yan
Shao, Shuo
Zheng, Haolun
Liu, Zhihao
Li, Jinfeng
Qin, Zhizhen
Chen, Yuefeng
Chu, Zhixuan
Qin, Zhan
Ren, Kui
author_facet Chen, Tailun
He, Yu
Wang, Yan
Shao, Shuo
Zheng, Haolun
Liu, Zhihao
Li, Jinfeng
Qin, Zhizhen
Chen, Yuefeng
Chu, Zhixuan
Qin, Zhan
Ren, Kui
contents Retrieval-Augmented Generation (RAG) systems enhance LLMs with external knowledge but introduce a critical attack surface: corpus poisoning. While recent studies have demonstrated the potential of such attacks, they typically rely on impractical assumptions, such as white-box access or known user queries, thereby underestimating the difficulty of real-world exploitation. In this paper, we bridge this gap by proposing MIRAGE, a novel multi-stage poisoning pipeline designed for strict black-box and query-agnostic environments. Operating on surrogate model feedback, MIRAGE functions as an automated optimization framework that integrates three key mechanisms: it utilizes persona-driven query synthesis to approximate latent user search distributions, employs semantic anchoring to imperceptibly embed these intents for high retrieval visibility, and leverages an adversarial variant of Test-Time Preference Optimization (TPO) to maximize persuasion. To rigorously evaluate this threat, we construct a new benchmark derived from three long-form, domain-specific datasets. Extensive experiments demonstrate that MIRAGE significantly outperforms existing baselines in both attack efficacy and stealthiness, exhibiting remarkable transferability across diverse retriever-LLM configurations and highlighting the urgent need for robust defense strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08289
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
Chen, Tailun
He, Yu
Wang, Yan
Shao, Shuo
Zheng, Haolun
Liu, Zhihao
Li, Jinfeng
Qin, Zhizhen
Chen, Yuefeng
Chu, Zhixuan
Qin, Zhan
Ren, Kui
Cryptography and Security
Retrieval-Augmented Generation (RAG) systems enhance LLMs with external knowledge but introduce a critical attack surface: corpus poisoning. While recent studies have demonstrated the potential of such attacks, they typically rely on impractical assumptions, such as white-box access or known user queries, thereby underestimating the difficulty of real-world exploitation. In this paper, we bridge this gap by proposing MIRAGE, a novel multi-stage poisoning pipeline designed for strict black-box and query-agnostic environments. Operating on surrogate model feedback, MIRAGE functions as an automated optimization framework that integrates three key mechanisms: it utilizes persona-driven query synthesis to approximate latent user search distributions, employs semantic anchoring to imperceptibly embed these intents for high retrieval visibility, and leverages an adversarial variant of Test-Time Preference Optimization (TPO) to maximize persuasion. To rigorously evaluate this threat, we construct a new benchmark derived from three long-form, domain-specific datasets. Extensive experiments demonstrate that MIRAGE significantly outperforms existing baselines in both attack efficacy and stealthiness, exhibiting remarkable transferability across diverse retriever-LLM configurations and highlighting the urgent need for robust defense strategies.
title MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
topic Cryptography and Security
url https://arxiv.org/abs/2512.08289