Certifiably Robust RAG against Retrieval Corruption

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Chong, Wu, Tong, Zhong, Zexuan, Wagner, David, Chen, Danqi, Mittal, Prateek
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911558312919040
author Xiang, Chong
Wu, Tong
Zhong, Zexuan
Wagner, David
Chen, Danqi
Mittal, Prateek
author_facet Xiang, Chong
Wu, Tong
Zhong, Zexuan
Wagner, David
Chen, Danqi
Mittal, Prateek
contents Retrieval-augmented generation (RAG) is susceptible to retrieval corruption attacks, where malicious passages injected into retrieval results can lead to inaccurate model responses. We propose RobustRAG, the first defense framework with certifiable robustness against retrieval corruption attacks. The key insight of RobustRAG is an isolate-then-aggregate strategy: we isolate passages into disjoint groups, generate LLM responses based on the concatenated passages from each isolated group, and then securely aggregate these responses for a robust output. To instantiate RobustRAG, we design keyword-based and decoding-based algorithms for securely aggregating unstructured text responses. Notably, RobustRAG achieves certifiable robustness: for certain queries in our evaluation datasets, we can formally certify non-trivial lower bounds on response quality -- even against an adaptive attacker with full knowledge of the defense and the ability to arbitrarily inject a bounded number of malicious passages. We evaluate RobustRAG on the tasks of open-domain question-answering and free-form long text generation and demonstrate its effectiveness across three datasets and three LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15556
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Certifiably Robust RAG against Retrieval Corruption
Xiang, Chong
Wu, Tong
Zhong, Zexuan
Wagner, David
Chen, Danqi
Mittal, Prateek
Machine Learning
Computation and Language
Cryptography and Security
Retrieval-augmented generation (RAG) is susceptible to retrieval corruption attacks, where malicious passages injected into retrieval results can lead to inaccurate model responses. We propose RobustRAG, the first defense framework with certifiable robustness against retrieval corruption attacks. The key insight of RobustRAG is an isolate-then-aggregate strategy: we isolate passages into disjoint groups, generate LLM responses based on the concatenated passages from each isolated group, and then securely aggregate these responses for a robust output. To instantiate RobustRAG, we design keyword-based and decoding-based algorithms for securely aggregating unstructured text responses. Notably, RobustRAG achieves certifiable robustness: for certain queries in our evaluation datasets, we can formally certify non-trivial lower bounds on response quality -- even against an adaptive attacker with full knowledge of the defense and the ability to arbitrarily inject a bounded number of malicious passages. We evaluate RobustRAG on the tasks of open-domain question-answering and free-form long text generation and demonstrate its effectiveness across three datasets and three LLMs.
title Certifiably Robust RAG against Retrieval Corruption
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2405.15556