MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ha, Hyeonjeong, Zhan, Qiusi, Kim, Jeonghwan, Bralios, Dimitrios, Sanniboina, Saikrishna, Peng, Nanyun, Chang, Kai-Wei, Kang, Daniel, Ji, Heng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916051330007040
author Ha, Hyeonjeong
Zhan, Qiusi
Kim, Jeonghwan
Bralios, Dimitrios
Sanniboina, Saikrishna
Peng, Nanyun
Chang, Kai-Wei
Kang, Daniel
Ji, Heng
author_facet Ha, Hyeonjeong
Zhan, Qiusi
Kim, Jeonghwan
Bralios, Dimitrios
Sanniboina, Saikrishna
Peng, Nanyun
Chang, Kai-Wei
Kang, Daniel
Ji, Heng
contents Retrieval-augmented generation (RAG) has become a common practice in multimodal large language models (MLLM) to enhance factual grounding and reduce hallucination. Yet, its reliance on retrieval exposes MLLMs to knowledge poisoning attacks, in which adversaries deliberately inject malicious multimodal content into external knowledge bases to steer models toward generating incorrect or even harmful responses. We present MM-PoisonRAG, a framework to systematically study the vulnerability of multimodal RAG under knowledge poisoning. Specifically, we design two novel attack strategies: Localized Poisoning Attack (LPA), which implants targeted, query-specific multimodal misinformation to manipulate outputs toward attacker-controlled responses, and Globalized Poisoning Attack (GPA), which uses a single, untargeted adversarial injection to broadly corrupt reasoning and collapse generation quality across all queries. Extensive experiments on diverse tasks, multimodal RAG components, and attacker access levels reveal severe vulnerabilities: LPA achieves up to 56% attack success rate even under restricted access, and transfers effectively across four different retrievers without re-optimizing the adversaries. GPA completely disrupts model generation to 0% accuracy with just one poisoned content. Moreover, both LPA and GPA bypass existing defenses, underscoring the fragility of multimodal RAG and establishing MM-PoisonRAG as a foundation for future research on securing RAG frameworks against multimodal knowledge poisoning.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks
Ha, Hyeonjeong
Zhan, Qiusi
Kim, Jeonghwan
Bralios, Dimitrios
Sanniboina, Saikrishna
Peng, Nanyun
Chang, Kai-Wei
Kang, Daniel
Ji, Heng
Machine Learning
Artificial Intelligence
Cryptography and Security
Computer Vision and Pattern Recognition
Retrieval-augmented generation (RAG) has become a common practice in multimodal large language models (MLLM) to enhance factual grounding and reduce hallucination. Yet, its reliance on retrieval exposes MLLMs to knowledge poisoning attacks, in which adversaries deliberately inject malicious multimodal content into external knowledge bases to steer models toward generating incorrect or even harmful responses. We present MM-PoisonRAG, a framework to systematically study the vulnerability of multimodal RAG under knowledge poisoning. Specifically, we design two novel attack strategies: Localized Poisoning Attack (LPA), which implants targeted, query-specific multimodal misinformation to manipulate outputs toward attacker-controlled responses, and Globalized Poisoning Attack (GPA), which uses a single, untargeted adversarial injection to broadly corrupt reasoning and collapse generation quality across all queries. Extensive experiments on diverse tasks, multimodal RAG components, and attacker access levels reveal severe vulnerabilities: LPA achieves up to 56% attack success rate even under restricted access, and transfers effectively across four different retrievers without re-optimizing the adversaries. GPA completely disrupts model generation to 0% accuracy with just one poisoned content. Moreover, both LPA and GPA bypass existing defenses, underscoring the fragility of multimodal RAG and establishing MM-PoisonRAG as a foundation for future research on securing RAG frameworks against multimodal knowledge poisoning.
title MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks
topic Machine Learning
Artificial Intelligence
Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.17832