HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yucheng, Li, Qinfeng, Du, Tianyu, Zhang, Xuhong, Zhao, Xinkui, Feng, Zhengwen, Yin, Jianwei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929568524271616
author Zhang, Yucheng
Li, Qinfeng
Du, Tianyu
Zhang, Xuhong
Zhao, Xinkui
Feng, Zhengwen
Yin, Jianwei
author_facet Zhang, Yucheng
Li, Qinfeng
Du, Tianyu
Zhang, Xuhong
Zhao, Xinkui
Feng, Zhengwen
Yin, Jianwei
contents Retrieval-Augmented Generation (RAG) systems enhance large language models (LLMs) by integrating external knowledge, making them adaptable and cost-effective for various applications. However, the growing reliance on these systems also introduces potential security risks. In this work, we reveal a novel vulnerability, the retrieval prompt hijack attack (HijackRAG), which enables attackers to manipulate the retrieval mechanisms of RAG systems by injecting malicious texts into the knowledge database. When the RAG system encounters target questions, it generates the attacker's pre-determined answers instead of the correct ones, undermining the integrity and trustworthiness of the system. We formalize HijackRAG as an optimization problem and propose both black-box and white-box attack strategies tailored to different levels of the attacker's knowledge. Extensive experiments on multiple benchmark datasets show that HijackRAG consistently achieves high attack success rates, outperforming existing baseline attacks. Furthermore, we demonstrate that the attack is transferable across different retriever models, underscoring the widespread risk it poses to RAG systems. Lastly, our exploration of various defense mechanisms reveals that they are insufficient to counter HijackRAG, emphasizing the urgent need for more robust security measures to protect RAG systems in real-world deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22832
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
Zhang, Yucheng
Li, Qinfeng
Du, Tianyu
Zhang, Xuhong
Zhao, Xinkui
Feng, Zhengwen
Yin, Jianwei
Cryptography and Security
Artificial Intelligence
Information Retrieval
Retrieval-Augmented Generation (RAG) systems enhance large language models (LLMs) by integrating external knowledge, making them adaptable and cost-effective for various applications. However, the growing reliance on these systems also introduces potential security risks. In this work, we reveal a novel vulnerability, the retrieval prompt hijack attack (HijackRAG), which enables attackers to manipulate the retrieval mechanisms of RAG systems by injecting malicious texts into the knowledge database. When the RAG system encounters target questions, it generates the attacker's pre-determined answers instead of the correct ones, undermining the integrity and trustworthiness of the system. We formalize HijackRAG as an optimization problem and propose both black-box and white-box attack strategies tailored to different levels of the attacker's knowledge. Extensive experiments on multiple benchmark datasets show that HijackRAG consistently achieves high attack success rates, outperforming existing baseline attacks. Furthermore, we demonstrate that the attack is transferable across different retriever models, underscoring the widespread risk it poses to RAG systems. Lastly, our exploration of various defense mechanisms reveals that they are insufficient to counter HijackRAG, emphasizing the urgent need for more robust security measures to protect RAG systems in real-world deployments.
title HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
topic Cryptography and Security
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2410.22832