CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kan, Zhehan, Zhang, Ce, Liao, Zihan, Tian, Yapeng, Yang, Wenming, Xiao, Junyuan, Li, Xu, Jiang, Dongmei, Wang, Yaowei, Liao, Qingmin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929597941022720
author Kan, Zhehan
Zhang, Ce
Liao, Zihan
Tian, Yapeng
Yang, Wenming
Xiao, Junyuan
Li, Xu
Jiang, Dongmei
Wang, Yaowei
Liao, Qingmin
author_facet Kan, Zhehan
Zhang, Ce
Liao, Zihan
Tian, Yapeng
Yang, Wenming
Xiao, Junyuan
Li, Xu
Jiang, Dongmei
Wang, Yaowei
Liao, Qingmin
contents Large Vision-Language Model (LVLM) systems have demonstrated impressive vision-language reasoning capabilities but suffer from pervasive and severe hallucination issues, posing significant risks in critical domains such as healthcare and autonomous systems. Despite previous efforts to mitigate hallucinations, a persistent issue remains: visual defect from vision-language misalignment, creating a bottleneck in visual processing capacity. To address this challenge, we develop Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs (CATCH), based on the Information Bottleneck theory. CATCH introduces Complementary Visual Decoupling (CVD) for visual information separation, Non-Visual Screening (NVS) for hallucination detection, and Adaptive Token-level Contrastive Decoding (ATCD) for hallucination mitigation. CATCH addresses issues related to visual defects that cause diminished fine-grained feature perception and cumulative hallucinations in open-ended scenarios. It is applicable to various visual question-answering tasks without requiring any specific data or prior knowledge, and generalizes robustly to new tasks without additional training, opening new possibilities for advancing LVLM in various challenging applications.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12713
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
Kan, Zhehan
Zhang, Ce
Liao, Zihan
Tian, Yapeng
Yang, Wenming
Xiao, Junyuan
Li, Xu
Jiang, Dongmei
Wang, Yaowei
Liao, Qingmin
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Model (LVLM) systems have demonstrated impressive vision-language reasoning capabilities but suffer from pervasive and severe hallucination issues, posing significant risks in critical domains such as healthcare and autonomous systems. Despite previous efforts to mitigate hallucinations, a persistent issue remains: visual defect from vision-language misalignment, creating a bottleneck in visual processing capacity. To address this challenge, we develop Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs (CATCH), based on the Information Bottleneck theory. CATCH introduces Complementary Visual Decoupling (CVD) for visual information separation, Non-Visual Screening (NVS) for hallucination detection, and Adaptive Token-level Contrastive Decoding (ATCD) for hallucination mitigation. CATCH addresses issues related to visual defects that cause diminished fine-grained feature perception and cumulative hallucinations in open-ended scenarios. It is applicable to various visual question-answering tasks without requiring any specific data or prior knowledge, and generalizes robustly to new tasks without additional training, opening new possibilities for advancing LVLM in various challenging applications.
title CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.12713