Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
Fuente:
arXiv
Salvato in:
| Autori principali: | Shafran, Avital, Schuster, Roei, Shmatikov, Vitaly |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rerouting LLM Routers
di: Shafran, Avital, et al.
Pubblicazione: (2025)
di: Shafran, Avital, et al.
Pubblicazione: (2025)
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
di: Zhang, Collin, et al.
Pubblicazione: (2024)
di: Zhang, Collin, et al.
Pubblicazione: (2024)
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
di: Mori, Junki, et al.
Pubblicazione: (2025)
di: Mori, Junki, et al.
Pubblicazione: (2025)
Multi-Agent Systems Execute Arbitrary Malicious Code
di: Triedman, Harold, et al.
Pubblicazione: (2025)
di: Triedman, Harold, et al.
Pubblicazione: (2025)
Laundering AI Authority with Adversarial Examples
di: Zhang, Jie, et al.
Pubblicazione: (2026)
di: Zhang, Jie, et al.
Pubblicazione: (2026)
Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
di: Jha, Rishi, et al.
Pubblicazione: (2025)
di: Jha, Rishi, et al.
Pubblicazione: (2025)
Certifiably Robust RAG against Retrieval Corruption
di: Xiang, Chong, et al.
Pubblicazione: (2024)
di: Xiang, Chong, et al.
Pubblicazione: (2024)
Universal Zero-shot Embedding Inversion
di: Zhang, Collin, et al.
Pubblicazione: (2025)
di: Zhang, Collin, et al.
Pubblicazione: (2025)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
di: Chaudhari, Harsh, et al.
Pubblicazione: (2024)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2024)
Beyond Labeling Oracles: What does it mean to steal ML models?
di: Shafran, Avital, et al.
Pubblicazione: (2023)
di: Shafran, Avital, et al.
Pubblicazione: (2023)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
di: Wang, Linlin, et al.
Pubblicazione: (2025)
di: Wang, Linlin, et al.
Pubblicazione: (2025)
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
di: Jha, Rishi, et al.
Pubblicazione: (2026)
di: Jha, Rishi, et al.
Pubblicazione: (2026)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
di: Hu, Xiao, et al.
Pubblicazione: (2025)
di: Hu, Xiao, et al.
Pubblicazione: (2025)
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
di: Xue, Jiaqi, et al.
Pubblicazione: (2024)
di: Xue, Jiaqi, et al.
Pubblicazione: (2024)
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?
di: Wang, Shang, et al.
Pubblicazione: (2024)
di: Wang, Shang, et al.
Pubblicazione: (2024)
Adversarial Illusions in Multi-Modal Embeddings
di: Zhang, Tingwei, et al.
Pubblicazione: (2023)
di: Zhang, Tingwei, et al.
Pubblicazione: (2023)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
di: Zhao, Tianzhe, et al.
Pubblicazione: (2025)
di: Zhao, Tianzhe, et al.
Pubblicazione: (2025)
Enhancing Robustness of AI Offensive Code Generators via Data Augmentation
di: Improta, Cristina, et al.
Pubblicazione: (2023)
di: Improta, Cristina, et al.
Pubblicazione: (2023)
Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning
di: Korn, Samuel
Pubblicazione: (2026)
di: Korn, Samuel
Pubblicazione: (2026)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
di: Chen, Yen-Shan, et al.
Pubblicazione: (2025)
di: Chen, Yen-Shan, et al.
Pubblicazione: (2025)
LLMGuard: Guarding Against Unsafe LLM Behavior
di: Goyal, Shubh, et al.
Pubblicazione: (2024)
di: Goyal, Shubh, et al.
Pubblicazione: (2024)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
di: Qi, Zhenting, et al.
Pubblicazione: (2024)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
di: Cheng, Pengzhou, et al.
Pubblicazione: (2024)
di: Cheng, Pengzhou, et al.
Pubblicazione: (2024)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
di: Hines, Keegan, et al.
Pubblicazione: (2024)
di: Hines, Keegan, et al.
Pubblicazione: (2024)
Composite Backdoor Attacks Against Large Language Models
di: Huang, Hai, et al.
Pubblicazione: (2023)
di: Huang, Hai, et al.
Pubblicazione: (2023)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
di: Brown, Hannah, et al.
Pubblicazione: (2024)
di: Brown, Hannah, et al.
Pubblicazione: (2024)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
di: Zhou, Ying, et al.
Pubblicazione: (2024)
di: Zhou, Ying, et al.
Pubblicazione: (2024)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
di: Zou, Wei, et al.
Pubblicazione: (2024)
di: Zou, Wei, et al.
Pubblicazione: (2024)
AutoBnB-RAG: Enhancing Multi-Agent Incident Response with Retrieval-Augmented Generation
di: Liu, Zefang, et al.
Pubblicazione: (2025)
di: Liu, Zefang, et al.
Pubblicazione: (2025)
Learned-Database Systems Security
di: Schuster, Roei, et al.
Pubblicazione: (2022)
di: Schuster, Roei, et al.
Pubblicazione: (2022)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
di: Yang, Ziqing, et al.
Pubblicazione: (2024)
di: Yang, Ziqing, et al.
Pubblicazione: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
di: Kim, Heegyu, et al.
Pubblicazione: (2024)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization
di: Wang, Lionel Z., et al.
Pubblicazione: (2026)
di: Wang, Lionel Z., et al.
Pubblicazione: (2026)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against
di: Enkhbayar, Tsogt-Ochir
Pubblicazione: (2025)
di: Enkhbayar, Tsogt-Ochir
Pubblicazione: (2025)
Self-interpreting Adversarial Images
di: Zhang, Tingwei, et al.
Pubblicazione: (2024)
di: Zhang, Tingwei, et al.
Pubblicazione: (2024)
Differential Degradation Vulnerabilities in Censorship Circumvention Systems
di: Sun, Zhen, et al.
Pubblicazione: (2024)
di: Sun, Zhen, et al.
Pubblicazione: (2024)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
di: Chen, Shuo, et al.
Pubblicazione: (2024)
di: Chen, Shuo, et al.
Pubblicazione: (2024)
SPML: A DSL for Defending Language Models Against Prompt Attacks
di: Sharma, Reshabh K, et al.
Pubblicazione: (2024)
di: Sharma, Reshabh K, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Rerouting LLM Routers
di: Shafran, Avital, et al.
Pubblicazione: (2025) -
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
di: Zhang, Collin, et al.
Pubblicazione: (2024) -
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
di: Mori, Junki, et al.
Pubblicazione: (2025) -
Multi-Agent Systems Execute Arbitrary Malicious Code
di: Triedman, Harold, et al.
Pubblicazione: (2025) -
Laundering AI Authority with Adversarial Examples
di: Zhang, Jie, et al.
Pubblicazione: (2026)