Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yen-Shan, Huang, Sian-Yao, Yang, Cheng-Lin, Chen, Yun-Nung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
by: Wang, Linlin, et al.
Published: (2025)
by: Wang, Linlin, et al.
Published: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning
by: Korn, Samuel
Published: (2026)
by: Korn, Samuel
Published: (2026)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
by: Zhao, Tianzhe, et al.
Published: (2025)
by: Zhao, Tianzhe, et al.
Published: (2025)
Learning to Poison Large Language Models for Downstream Manipulation
by: Zhou, Xiangyu, et al.
Published: (2024)
by: Zhou, Xiangyu, et al.
Published: (2024)
Certifiably Robust RAG against Retrieval Corruption
by: Xiang, Chong, et al.
Published: (2024)
by: Xiang, Chong, et al.
Published: (2024)
Transferable Availability Poisoning Attacks
by: Liu, Yiyong, et al.
Published: (2023)
by: Liu, Yiyong, et al.
Published: (2023)
Poisoning Attacks to Local Differential Privacy Protocols for Trajectory Data
by: Hsu, I-Jung, et al.
Published: (2025)
by: Hsu, I-Jung, et al.
Published: (2025)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
by: Hu, Haoyang, et al.
Published: (2026)
by: Hu, Haoyang, et al.
Published: (2026)
Fingerprint Vector: Enabling Scalable and Efficient Model Fingerprint Transfer via Vector Addition
by: Xu, Zhenhua, et al.
Published: (2024)
by: Xu, Zhenhua, et al.
Published: (2024)
Transferable Embedding Inversion Attack: Uncovering Privacy Risks in Text Embeddings without Model Queries
by: Huang, Yu-Hsiang, et al.
Published: (2024)
by: Huang, Yu-Hsiang, et al.
Published: (2024)
Provable Watermarking for Data Poisoning Attacks
by: Zhu, Yifan, et al.
Published: (2025)
by: Zhu, Yifan, et al.
Published: (2025)
Retrieval-Augmented Review Generation for Poisoning Recommender Systems
by: Yang, Shiyi, et al.
Published: (2025)
by: Yang, Shiyi, et al.
Published: (2025)
Denial-of-Service Poisoning Attacks against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
by: Hu, Hanjiang, et al.
Published: (2025)
by: Hu, Hanjiang, et al.
Published: (2025)
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
by: Chen, Bocheng, et al.
Published: (2024)
by: Chen, Bocheng, et al.
Published: (2024)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
by: Baumgärtner, Tim, et al.
Published: (2024)
by: Baumgärtner, Tim, et al.
Published: (2024)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
by: Zou, Wei, et al.
Published: (2024)
by: Zou, Wei, et al.
Published: (2024)
From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse
by: Jhuma, Rabeya Amin, et al.
Published: (2025)
by: Jhuma, Rabeya Amin, et al.
Published: (2025)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
by: Afane, Mohamed, et al.
Published: (2025)
by: Afane, Mohamed, et al.
Published: (2025)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
by: Yao, Duanyi, et al.
Published: (2026)
by: Yao, Duanyi, et al.
Published: (2026)
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
by: Mori, Junki, et al.
Published: (2025)
by: Mori, Junki, et al.
Published: (2025)
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
by: Shafran, Avital, et al.
Published: (2024)
by: Shafran, Avital, et al.
Published: (2024)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
by: Wu, Yu-Hang, et al.
Published: (2025)
by: Wu, Yu-Hang, et al.
Published: (2025)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Adversarial Data Poisoning for Fake News Detection: How to Make a Model Misclassify a Target News without Modifying It
by: Siciliano, Federico, et al.
Published: (2023)
by: Siciliano, Federico, et al.
Published: (2023)
On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic Writing
by: Liu, Zeyan, et al.
Published: (2023)
by: Liu, Zeyan, et al.
Published: (2023)
Universal Jailbreak Backdoors from Poisoned Human Feedback
by: Rando, Javier, et al.
Published: (2023)
by: Rando, Javier, et al.
Published: (2023)
RADAR: Defending RAG Dynamically against Retrieval Corruption
by: Chen, Ziyuan, et al.
Published: (2026)
by: Chen, Ziyuan, et al.
Published: (2026)
Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training
by: Duan, Wenjing, et al.
Published: (2026)
by: Duan, Wenjing, et al.
Published: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
by: Wang, Tianchun, et al.
Published: (2024)
by: Wang, Tianchun, et al.
Published: (2024)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
by: He, Haorui, et al.
Published: (2025)
by: He, Haorui, et al.
Published: (2025)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
by: Chaudhari, Harsh, et al.
Published: (2026)
by: Chaudhari, Harsh, et al.
Published: (2026)
Interpretable LLM Guardrails via Sparse Representation Steering
by: He, Zeqing, et al.
Published: (2025)
by: He, Zeqing, et al.
Published: (2025)
Similar Items
-
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
by: Chen, Yen-Shan, et al.
Published: (2026) -
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
by: Wang, Linlin, et al.
Published: (2025) -
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026) -
Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning
by: Korn, Samuel
Published: (2026) -
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
by: Zhao, Tianzhe, et al.
Published: (2025)