Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yining, Zhang, Mi, Sun, Junjie, Wang, Chenyue, Yang, Min, Xue, Hui, Tao, Jialing, Duan, Ranjie, Liu, Jiexi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910801771626496
author Wang, Yining
Zhang, Mi
Sun, Junjie
Wang, Chenyue
Yang, Min
Xue, Hui
Tao, Jialing
Duan, Ranjie
Liu, Jiexi
author_facet Wang, Yining
Zhang, Mi
Sun, Junjie
Wang, Chenyue
Yang, Min
Xue, Hui
Tao, Jialing
Duan, Ranjie
Liu, Jiexi
contents Fusing visual understanding into language generation, Multi-modal Large Language Models (MLLMs) are revolutionizing visual-language applications. Yet, these models are often plagued by the hallucination problem, which involves generating inaccurate objects, attributes, and relationships that do not match the visual content. In this work, we delve into the internal attention mechanisms of MLLMs to reveal the underlying causes of hallucination, exposing the inherent vulnerabilities in the instruction-tuning process. We propose a novel hallucination attack against MLLMs that exploits attention sink behaviors to trigger hallucinated content with minimal image-text relevance, posing a significant threat to critical downstream applications. Distinguished from previous adversarial methods that rely on fixed patterns, our approach generates dynamic, effective, and highly transferable visual adversarial inputs, without sacrificing the quality of model responses. Comprehensive experiments on 6 prominent MLLMs demonstrate the efficacy of our attack in compromising black-box MLLMs even with extensive mitigating mechanisms, as well as the promising results against cutting-edge commercial APIs, such as GPT-4o and Gemini 1.5. Our code is available at https://huggingface.co/RachelHGF/Mirage-in-the-Eyes.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
Wang, Yining
Zhang, Mi
Sun, Junjie
Wang, Chenyue
Yang, Min
Xue, Hui
Tao, Jialing
Duan, Ranjie
Liu, Jiexi
Machine Learning
Cryptography and Security
Computer Vision and Pattern Recognition
Fusing visual understanding into language generation, Multi-modal Large Language Models (MLLMs) are revolutionizing visual-language applications. Yet, these models are often plagued by the hallucination problem, which involves generating inaccurate objects, attributes, and relationships that do not match the visual content. In this work, we delve into the internal attention mechanisms of MLLMs to reveal the underlying causes of hallucination, exposing the inherent vulnerabilities in the instruction-tuning process. We propose a novel hallucination attack against MLLMs that exploits attention sink behaviors to trigger hallucinated content with minimal image-text relevance, posing a significant threat to critical downstream applications. Distinguished from previous adversarial methods that rely on fixed patterns, our approach generates dynamic, effective, and highly transferable visual adversarial inputs, without sacrificing the quality of model responses. Comprehensive experiments on 6 prominent MLLMs demonstrate the efficacy of our attack in compromising black-box MLLMs even with extensive mitigating mechanisms, as well as the promising results against cutting-edge commercial APIs, such as GPT-4o and Gemini 1.5. Our code is available at https://huggingface.co/RachelHGF/Mirage-in-the-Eyes.
title Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
topic Machine Learning
Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.15269