Certifiably Robust RAG against Retrieval Corruption
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Chong, Wu, Tong, Zhong, Zexuan, Wagner, David, Chen, Danqi, Mittal, Prateek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial Patches
by: Jacob, Dennis, et al.
Published: (2025)
by: Jacob, Dennis, et al.
Published: (2025)
Position: Towards Resilience Against Adversarial Examples
by: Dai, Sihui, et al.
Published: (2024)
by: Dai, Sihui, et al.
Published: (2024)
RADAR: Defending RAG Dynamically against Retrieval Corruption
by: Chen, Ziyuan, et al.
Published: (2026)
by: Chen, Ziyuan, et al.
Published: (2026)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
by: Lou, Qian, et al.
Published: (2024)
by: Lou, Qian, et al.
Published: (2024)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Private Fine-tuning of Large Language Models with Zeroth-order Optimization
by: Tang, Xinyu, et al.
Published: (2024)
by: Tang, Xinyu, et al.
Published: (2024)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
CERT-ED: Certifiably Robust Text Classification for Edit Distance
by: Huang, Zhuoqun, et al.
Published: (2024)
by: Huang, Zhuoqun, et al.
Published: (2024)
AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness
by: Huang, Zhuoqun, et al.
Published: (2025)
by: Huang, Zhuoqun, et al.
Published: (2025)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
by: Wang, Linlin, et al.
Published: (2025)
by: Wang, Linlin, et al.
Published: (2025)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
by: Shafran, Avital, et al.
Published: (2024)
by: Shafran, Avital, et al.
Published: (2024)
Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG)
by: Mori, Junki, et al.
Published: (2025)
by: Mori, Junki, et al.
Published: (2025)
PatchCURE: Improving Certifiable Robustness, Model Utility, and Computation Efficiency of Adversarial Patch Defenses
by: Xiang, Chong, et al.
Published: (2023)
by: Xiang, Chong, et al.
Published: (2023)
ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
by: Shen, Zeyu, et al.
Published: (2025)
by: Shen, Zeyu, et al.
Published: (2025)
Model Provenance Testing for Large Language Models
by: Nikolic, Ivica, et al.
Published: (2025)
by: Nikolic, Ivica, et al.
Published: (2025)
Teach LLMs to Phish: Stealing Private Information from Language Models
by: Panda, Ashwinee, et al.
Published: (2024)
by: Panda, Ashwinee, et al.
Published: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
by: Hu, Hanjiang, et al.
Published: (2025)
by: Hu, Hanjiang, et al.
Published: (2025)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
by: Zhang, Fangyuan, et al.
Published: (2024)
by: Zhang, Fangyuan, et al.
Published: (2024)
Collective Certified Robustness against Graph Injection Attacks
by: Lai, Yuni, et al.
Published: (2024)
by: Lai, Yuni, et al.
Published: (2024)
Detecting Pretraining Data from Large Language Models
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
by: Chen, Yen-Shan, et al.
Published: (2025)
by: Chen, Yen-Shan, et al.
Published: (2025)
Benchmarking Large Multimodal Models against Common Corruptions
by: Zhang, Jiawei, et al.
Published: (2024)
by: Zhang, Jiawei, et al.
Published: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
by: Chen, Taiye, et al.
Published: (2025)
by: Chen, Taiye, et al.
Published: (2025)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
by: Zou, Wei, et al.
Published: (2024)
by: Zou, Wei, et al.
Published: (2024)
Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning
by: Korn, Samuel
Published: (2026)
by: Korn, Samuel
Published: (2026)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
by: Hu, Xiao, et al.
Published: (2025)
by: Hu, Xiao, et al.
Published: (2025)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
by: Zhu, Kaijie, et al.
Published: (2023)
by: Zhu, Kaijie, et al.
Published: (2023)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
by: Fu, Wenjie, et al.
Published: (2023)
by: Fu, Wenjie, et al.
Published: (2023)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
by: Geng, Runpeng, et al.
Published: (2025)
by: Geng, Runpeng, et al.
Published: (2025)
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
RED QUEEN: Safeguarding Large Language Models against Concealed Multi-Turn Jailbreaking
by: Jiang, Yifan, et al.
Published: (2024)
by: Jiang, Yifan, et al.
Published: (2024)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
DocMIA: Document-Level Membership Inference Attacks against DocVQA Models
by: Nguyen, Khanh, et al.
Published: (2025)
by: Nguyen, Khanh, et al.
Published: (2025)
Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
by: Dang, Trung Cuong, et al.
Published: (2025)
by: Dang, Trung Cuong, et al.
Published: (2025)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
by: He, Jiaming, et al.
Published: (2024)
by: He, Jiaming, et al.
Published: (2024)
Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection
by: Kimura, Subaru, et al.
Published: (2024)
by: Kimura, Subaru, et al.
Published: (2024)
Similar Items
-
PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial Patches
by: Jacob, Dennis, et al.
Published: (2025) -
Position: Towards Resilience Against Adversarial Examples
by: Dai, Sihui, et al.
Published: (2024) -
RADAR: Defending RAG Dynamically against Retrieval Corruption
by: Chen, Ziyuan, et al.
Published: (2026) -
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
by: Lou, Qian, et al.
Published: (2024) -
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)