DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Yudong, Xie, Ruobing, Sun, Xingwu, Huang, Yiqing, Chen, Jiansheng, Kang, Zhanhui, Wang, Di, Wang, Yu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911084124831744
author Zhang, Yudong
Xie, Ruobing
Sun, Xingwu
Huang, Yiqing
Chen, Jiansheng
Kang, Zhanhui
Wang, Di
Wang, Yu
author_facet Zhang, Yudong
Xie, Ruobing
Sun, Xingwu
Huang, Yiqing
Chen, Jiansheng
Kang, Zhanhui
Wang, Di
Wang, Yu
contents Large vision-language models (LVLMs) have demonstrated exceptional performance on complex multimodal tasks. However, they continue to suffer from significant hallucination issues, including object, attribute, and relational hallucinations. To accurately detect these hallucinations, we investigated the variations in cross-modal attention patterns between hallucination and non-hallucination states. Leveraging these distinctions, we developed a lightweight detector capable of identifying hallucinations. Our proposed method, Detecting Hallucinations by Cross-modal Attention Patterns (DHCP), is straightforward and does not require additional LVLM training or extra LVLM inference steps. Experimental results show that DHCP achieves remarkable performance in hallucination detection. By offering novel insights into the identification and analysis of hallucinations in LVLMs, DHCP contributes to advancing the reliability and trustworthiness of these models. The code is available at https://github.com/btzyd/DHCP.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18659
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
Zhang, Yudong
Xie, Ruobing
Sun, Xingwu
Huang, Yiqing
Chen, Jiansheng
Kang, Zhanhui
Wang, Di
Wang, Yu
Computer Vision and Pattern Recognition
Artificial Intelligence
Large vision-language models (LVLMs) have demonstrated exceptional performance on complex multimodal tasks. However, they continue to suffer from significant hallucination issues, including object, attribute, and relational hallucinations. To accurately detect these hallucinations, we investigated the variations in cross-modal attention patterns between hallucination and non-hallucination states. Leveraging these distinctions, we developed a lightweight detector capable of identifying hallucinations. Our proposed method, Detecting Hallucinations by Cross-modal Attention Patterns (DHCP), is straightforward and does not require additional LVLM training or extra LVLM inference steps. Experimental results show that DHCP achieves remarkable performance in hallucination detection. By offering novel insights into the identification and analysis of hallucinations in LVLMs, DHCP contributes to advancing the reliability and trustworthiness of these models. The code is available at https://github.com/btzyd/DHCP.
title DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.18659