Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shafran, Avital, Schuster, Roei, Shmatikov, Vitaly
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910865797677056
author Shafran, Avital
Schuster, Roei
Shmatikov, Vitaly
author_facet Shafran, Avital
Schuster, Roei
Shmatikov, Vitaly
contents Retrieval-augmented generation (RAG) systems respond to queries by retrieving relevant documents from a knowledge database and applying an LLM to the retrieved documents. We demonstrate that RAG systems that operate on databases with untrusted content are vulnerable to denial-of-service attacks we call jamming. An adversary can add a single ``blocker'' document to the database that will be retrieved in response to a specific query and result in the RAG system not answering this query, ostensibly because it lacks relevant information or because the answer is unsafe. We describe and measure the efficacy of several methods for generating blocker documents, including a new method based on black-box optimization. Our method (1) does not rely on instruction injection, (2) does not require the adversary to know the embedding or LLM used by the target RAG system, and (3) does not employ an auxiliary LLM. We evaluate jamming attacks on several embeddings and LLMs and demonstrate that the existing safety metrics for LLMs do not capture their vulnerability to jamming. We then discuss defenses against blocker documents.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05870
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
Shafran, Avital
Schuster, Roei
Shmatikov, Vitaly
Cryptography and Security
Computation and Language
Machine Learning
Retrieval-augmented generation (RAG) systems respond to queries by retrieving relevant documents from a knowledge database and applying an LLM to the retrieved documents. We demonstrate that RAG systems that operate on databases with untrusted content are vulnerable to denial-of-service attacks we call jamming. An adversary can add a single ``blocker'' document to the database that will be retrieved in response to a specific query and result in the RAG system not answering this query, ostensibly because it lacks relevant information or because the answer is unsafe. We describe and measure the efficacy of several methods for generating blocker documents, including a new method based on black-box optimization. Our method (1) does not rely on instruction injection, (2) does not require the adversary to know the embedding or LLM used by the target RAG system, and (3) does not employ an auxiliary LLM. We evaluate jamming attacks on several embeddings and LLMs and demonstrate that the existing safety metrics for LLMs do not capture their vulnerability to jamming. We then discuss defenses against blocker documents.
title Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
topic Cryptography and Security
Computation and Language
Machine Learning
url https://arxiv.org/abs/2406.05870