Disabling Self-Correction in Retrieval-Augmented Generation via Stealthy Retriever Poisoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dai, Yanbo, Ji, Zhenlan, Li, Zongjie, Li, Kuan, Wang, Shuai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SEAL: Subspace-Anchored Watermarks for LLM Ownership
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Review Generation for Poisoning Recommender Systems
von: Yang, Shiyi, et al.
Veröffentlicht: (2025)
von: Yang, Shiyi, et al.
Veröffentlicht: (2025)
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
von: Wang, Liwen, et al.
Veröffentlicht: (2025)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
von: Naseh, Ali, et al.
Veröffentlicht: (2025)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
EAMET: Robust Massive Model Editing via Embedding Alignment Optimization
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
von: Dai, Yanbo, et al.
Veröffentlicht: (2025)
Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors
von: Yin, Rui, et al.
Veröffentlicht: (2026)
von: Yin, Rui, et al.
Veröffentlicht: (2026)
Differentially Private Retrieval-Augmented Generation
von: Tang, Tingting, et al.
Veröffentlicht: (2026)
von: Tang, Tingting, et al.
Veröffentlicht: (2026)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
von: Deng, Gelei, et al.
Veröffentlicht: (2024)
von: Deng, Gelei, et al.
Veröffentlicht: (2024)
SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems
von: Bodea, Andreea-Elena, et al.
Veröffentlicht: (2026)
von: Bodea, Andreea-Elena, et al.
Veröffentlicht: (2026)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
von: Wang, Linlin, et al.
Veröffentlicht: (2025)
von: Wang, Linlin, et al.
Veröffentlicht: (2025)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?
von: Wang, Shang, et al.
Veröffentlicht: (2024)
von: Wang, Shang, et al.
Veröffentlicht: (2024)
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
von: Zhang, Jiankun, et al.
Veröffentlicht: (2025)
von: Zhang, Jiankun, et al.
Veröffentlicht: (2025)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
SoK: Evaluating Jailbreak Guardrails for Large Language Models
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
von: Wang, Xunguang, et al.
Veröffentlicht: (2025)
Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training
von: Duan, Wenjing, et al.
Veröffentlicht: (2026)
von: Duan, Wenjing, et al.
Veröffentlicht: (2026)
AutoBnB-RAG: Enhancing Multi-Agent Incident Response with Retrieval-Augmented Generation
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
von: Koga, Tatsuki, et al.
Veröffentlicht: (2024)
von: Koga, Tatsuki, et al.
Veröffentlicht: (2024)
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
von: Shereen, Ezzeldin, et al.
Veröffentlicht: (2025)
von: Shereen, Ezzeldin, et al.
Veröffentlicht: (2025)
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
von: Mori, Junki, et al.
Veröffentlicht: (2025)
von: Mori, Junki, et al.
Veröffentlicht: (2025)
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2024)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
von: Yang, Guangyu, et al.
Veröffentlicht: (2025)
von: Yang, Guangyu, et al.
Veröffentlicht: (2025)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
von: Shafran, Avital, et al.
Veröffentlicht: (2024)
von: Shafran, Avital, et al.
Veröffentlicht: (2024)
MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
von: Chen, Tailun, et al.
Veröffentlicht: (2025)
von: Chen, Tailun, et al.
Veröffentlicht: (2025)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
A Decentralized Retrieval Augmented Generation System with Source Reliabilities Secured on Blockchain
von: Lu, Yining, et al.
Veröffentlicht: (2025)
von: Lu, Yining, et al.
Veröffentlicht: (2025)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
von: Geng, Runpeng, et al.
Veröffentlicht: (2025)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
von: Chen, Zhuo, et al.
Veröffentlicht: (2024)
Traceback of Poisoning Attacks to Retrieval-Augmented Generation
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
von: Zhang, Baolei, et al.
Veröffentlicht: (2025)
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
von: Liu, Yinuo, et al.
Veröffentlicht: (2025)
Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models
von: Gong, Yuyang, et al.
Veröffentlicht: (2025)
von: Gong, Yuyang, et al.
Veröffentlicht: (2025)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SEAL: Subspace-Anchored Watermarks for LLM Ownership
von: Dai, Yanbo, et al.
Veröffentlicht: (2025) -
Retrieval-Augmented Review Generation for Poisoning Recommender Systems
von: Yang, Shiyi, et al.
Veröffentlicht: (2025) -
STShield: Single-Token Sentinel for Real-Time Jailbreak Detection in Large Language Models
von: Wang, Xunguang, et al.
Veröffentlicht: (2025) -
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025) -
IP Leakage Attacks Targeting LLM-Based Multi-Agent Systems
von: Wang, Liwen, et al.
Veröffentlicht: (2025)