Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Linlin, Zhu, Tianqing, Qin, Laiqiao, Gao, Longxiang, Zhou, Wanlei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?
by: Wang, Shang, et al.
Published: (2024)
by: Wang, Shang, et al.
Published: (2024)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
by: Chen, Yen-Shan, et al.
Published: (2025)
by: Chen, Yen-Shan, et al.
Published: (2025)
Linkage on Security, Privacy and Fairness in Federated Learning: New Balances and New Perspectives
by: Wang, Linlin, et al.
Published: (2024)
by: Wang, Linlin, et al.
Published: (2024)
Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning
by: Korn, Samuel
Published: (2026)
by: Korn, Samuel
Published: (2026)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
by: Zhao, Tianzhe, et al.
Published: (2025)
by: Zhao, Tianzhe, et al.
Published: (2025)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
by: Sivapiromrat, Sanhanat, et al.
Published: (2025)
Certifiably Robust RAG against Retrieval Corruption
by: Xiang, Chong, et al.
Published: (2024)
by: Xiang, Chong, et al.
Published: (2024)
Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA
by: Qin, Laiqiao, et al.
Published: (2024)
by: Qin, Laiqiao, et al.
Published: (2024)
Learning to Poison Large Language Models for Downstream Manipulation
by: Zhou, Xiangyu, et al.
Published: (2024)
by: Zhou, Xiangyu, et al.
Published: (2024)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
by: Zou, Wei, et al.
Published: (2024)
by: Zou, Wei, et al.
Published: (2024)
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
by: Mori, Junki, et al.
Published: (2025)
by: Mori, Junki, et al.
Published: (2025)
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
by: Shafran, Avital, et al.
Published: (2024)
by: Shafran, Avital, et al.
Published: (2024)
Towards Efficient Target-Level Machine Unlearning Based on Essential Graph
by: Xu, Heng, et al.
Published: (2024)
by: Xu, Heng, et al.
Published: (2024)
Forgetting Similar Samples: Can Machine Unlearning Do it Better?
by: Xu, Heng, et al.
Published: (2026)
by: Xu, Heng, et al.
Published: (2026)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Defending Against Neural Network Model Inversion Attacks via Data Poisoning
by: Zhou, Shuai, et al.
Published: (2024)
by: Zhou, Shuai, et al.
Published: (2024)
Update Selective Parameters: Federated Machine Unlearning Based on Model Explanation
by: Xu, Heng, et al.
Published: (2024)
by: Xu, Heng, et al.
Published: (2024)
Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives
by: Chang, Wenhan, et al.
Published: (2025)
by: Chang, Wenhan, et al.
Published: (2025)
Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment
by: Shao, Zedian, et al.
Published: (2024)
by: Shao, Zedian, et al.
Published: (2024)
Zero-shot Class Unlearning via Layer-wise Relevance Analysis and Neuronal Path Perturbation
by: Chang, Wenhan, et al.
Published: (2024)
by: Chang, Wenhan, et al.
Published: (2024)
Knowledge Distillation in Federated Learning: a Survey on Long Lasting Challenges and New Solutions
by: Qin, Laiqiao, et al.
Published: (2024)
by: Qin, Laiqiao, et al.
Published: (2024)
Reinforcement Unlearning
by: Ye, Dayong, et al.
Published: (2023)
by: Ye, Dayong, et al.
Published: (2023)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
by: Park, Seong-Gyu, et al.
Published: (2026)
by: Park, Seong-Gyu, et al.
Published: (2026)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
by: Baumgärtner, Tim, et al.
Published: (2024)
by: Baumgärtner, Tim, et al.
Published: (2024)
When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
by: Hu, Hanjiang, et al.
Published: (2025)
by: Hu, Hanjiang, et al.
Published: (2025)
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse
by: Jhuma, Rabeya Amin, et al.
Published: (2025)
by: Jhuma, Rabeya Amin, et al.
Published: (2025)
Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection
by: Liu, Chi, et al.
Published: (2025)
by: Liu, Chi, et al.
Published: (2025)
Adversarial Data Poisoning for Fake News Detection: How to Make a Model Misclassify a Target News without Modifying It
by: Siciliano, Federico, et al.
Published: (2023)
by: Siciliano, Federico, et al.
Published: (2023)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
by: Liu, Bing, et al.
Published: (2026)
by: Liu, Bing, et al.
Published: (2026)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
Federated Learning with Blockchain-Enhanced Machine Unlearning: A Trustworthy Approach
by: Zuo, Xuhan, et al.
Published: (2024)
by: Zuo, Xuhan, et al.
Published: (2024)
When Fairness Meets Privacy: Exploring Privacy Threats in Fair Binary Classifiers via Membership Inference Attacks
by: Tian, Huan, et al.
Published: (2023)
by: Tian, Huan, et al.
Published: (2023)
How Does a Deep Learning Model Architecture Impact Its Privacy? A Comprehensive Study of Privacy Attacks on CNNs and Transformers
by: Zhang, Guangsheng, et al.
Published: (2022)
by: Zhang, Guangsheng, et al.
Published: (2022)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
by: Zhao, Xuandong, et al.
Published: (2024)
by: Zhao, Xuandong, et al.
Published: (2024)
Osmosis Distillation: Model Hijacking with the Fewest Samples
by: Shi, Yuchen, et al.
Published: (2026)
by: Shi, Yuchen, et al.
Published: (2026)
KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation
by: Chen, Qizhi, et al.
Published: (2026)
by: Chen, Qizhi, et al.
Published: (2026)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
by: Khattar, Vanshaj, et al.
Published: (2026)
by: Khattar, Vanshaj, et al.
Published: (2026)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
by: Li, Ran, et al.
Published: (2025)
by: Li, Ran, et al.
Published: (2025)
Similar Items
-
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?
by: Wang, Shang, et al.
Published: (2024) -
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
by: Chen, Yen-Shan, et al.
Published: (2025) -
Linkage on Security, Privacy and Fairness in Federated Learning: New Balances and New Perspectives
by: Wang, Linlin, et al.
Published: (2024) -
Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning
by: Korn, Samuel
Published: (2026) -
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
by: Zhao, Tianzhe, et al.
Published: (2025)