ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Junzhe, Zhang, Tianshu, Huang, Shiyu, Niu, Yuwei, Zhang, Linfeng, Wen, Lijie, Hu, Xuming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909400871993344
author Chen, Junzhe
Zhang, Tianshu
Huang, Shiyu
Niu, Yuwei
Zhang, Linfeng
Wen, Lijie
Hu, Xuming
author_facet Chen, Junzhe
Zhang, Tianshu
Huang, Shiyu
Niu, Yuwei
Zhang, Linfeng
Wen, Lijie
Hu, Xuming
contents Despite the recent breakthroughs achieved by Large Vision Language Models (LVLMs) in understanding and responding to complex visual-textual contexts, their inherent hallucination tendencies limit their practical application in real-world scenarios that demand high levels of precision. Existing methods typically either fine-tune the LVLMs using additional data, which incurs extra costs in manual annotation and computational resources or perform comparisons at the decoding stage, which may eliminate useful language priors for reasoning while introducing inference time overhead. Therefore, we propose ICT, a lightweight, training-free method that calculates an intervention direction to shift the model's focus towards different levels of visual information, enhancing its attention to high-level and fine-grained visual details. During the forward pass stage, the intervention is applied to the attention heads that encode the overall image information and the fine-grained object details, effectively mitigating the phenomenon of overly language priors, and thereby alleviating hallucinations. Extensive experiments demonstrate that ICT achieves strong performance with a small amount of data and generalizes well across different datasets and models. Our code will be public.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15268
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
Chen, Junzhe
Zhang, Tianshu
Huang, Shiyu
Niu, Yuwei
Zhang, Linfeng
Wen, Lijie
Hu, Xuming
Computer Vision and Pattern Recognition
Computation and Language
Despite the recent breakthroughs achieved by Large Vision Language Models (LVLMs) in understanding and responding to complex visual-textual contexts, their inherent hallucination tendencies limit their practical application in real-world scenarios that demand high levels of precision. Existing methods typically either fine-tune the LVLMs using additional data, which incurs extra costs in manual annotation and computational resources or perform comparisons at the decoding stage, which may eliminate useful language priors for reasoning while introducing inference time overhead. Therefore, we propose ICT, a lightweight, training-free method that calculates an intervention direction to shift the model's focus towards different levels of visual information, enhancing its attention to high-level and fine-grained visual details. During the forward pass stage, the intervention is applied to the attention heads that encode the overall image information and the fine-grained object details, effectively mitigating the phenomenon of overly language priors, and thereby alleviating hallucinations. Extensive experiments demonstrate that ICT achieves strong performance with a small amount of data and generalizes well across different datasets and models. Our code will be public.
title ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2411.15268