Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yen-Shan, Huang, Sian-Yao, Yang, Cheng-Lin, Chen, Yun-Nung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911305880829952
author Chen, Yen-Shan
Huang, Sian-Yao
Yang, Cheng-Lin
Chen, Yun-Nung
author_facet Chen, Yen-Shan
Huang, Sian-Yao
Yang, Cheng-Lin
Chen, Yun-Nung
contents Existing data poisoning attacks on retrieval-augmented generation (RAG) systems scale poorly because they require costly optimization of poisoned documents for each target phrase. We introduce Eyes-on-Me, a modular attack that decomposes an adversarial document into reusable Attention Attractors and Focus Regions. Attractors are optimized to direct attention to the Focus Region. Attackers can then insert semantic baits for the retriever or malicious instructions for the generator, adapting to new targets at near zero cost. This is achieved by steering a small subset of attention heads that we empirically identify as strongly correlated with attack success. Across 18 end-to-end RAG settings (3 datasets $\times$ 2 retrievers $\times$ 3 generators), Eyes-on-Me raises average attack success rates from 21.9 to 57.8 (+35.9 points, 2.6$\times$ over prior work). A single optimized attractor transfers to unseen black box retrievers and generators without retraining. Our findings establish a scalable paradigm for RAG data poisoning and show that modular, reusable components pose a practical threat to modern AI systems. They also reveal a strong link between attention concentration and model outputs, informing interpretability research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
Chen, Yen-Shan
Huang, Sian-Yao
Yang, Cheng-Lin
Chen, Yun-Nung
Machine Learning
Computation and Language
Cryptography and Security
Existing data poisoning attacks on retrieval-augmented generation (RAG) systems scale poorly because they require costly optimization of poisoned documents for each target phrase. We introduce Eyes-on-Me, a modular attack that decomposes an adversarial document into reusable Attention Attractors and Focus Regions. Attractors are optimized to direct attention to the Focus Region. Attackers can then insert semantic baits for the retriever or malicious instructions for the generator, adapting to new targets at near zero cost. This is achieved by steering a small subset of attention heads that we empirically identify as strongly correlated with attack success. Across 18 end-to-end RAG settings (3 datasets $\times$ 2 retrievers $\times$ 3 generators), Eyes-on-Me raises average attack success rates from 21.9 to 57.8 (+35.9 points, 2.6$\times$ over prior work). A single optimized attractor transfers to unseen black box retrievers and generators without retraining. Our findings establish a scalable paradigm for RAG data poisoning and show that modular, reusable components pose a practical threat to modern AI systems. They also reveal a strong link between attention concentration and model outputs, informing interpretability research.
title Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2510.00586