Understanding Multimodal Hallucination with Parameter-Free Representation Alignment

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yueqian, Liang, Jianxin, Wang, Yuxuan, Zhang, Huishuai, Zhao, Dongyan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929482673160192
author Wang, Yueqian
Liang, Jianxin
Wang, Yuxuan
Zhang, Huishuai
Zhao, Dongyan
author_facet Wang, Yueqian
Liang, Jianxin
Wang, Yuxuan
Zhang, Huishuai
Zhao, Dongyan
contents Hallucination is a common issue in Multimodal Large Language Models (MLLMs), yet the underlying principles remain poorly understood. In this paper, we investigate which components of MLLMs contribute to object hallucinations. To analyze image representations while completely avoiding the influence of all other factors other than the image representation itself, we propose a parametric-free representation alignment metric (Pfram) that can measure the similarities between any two representation systems without requiring additional training parameters. Notably, Pfram can also assess the alignment of a neural representation system with the human representation system, represented by ground-truth annotations of images. By evaluating the alignment with object annotations, we demonstrate that this metric shows strong and consistent correlations with object hallucination across a wide range of state-of-the-art MLLMs, spanning various model architectures and sizes. Furthermore, using this metric, we explore other key issues related to image representations in MLLMs, such as the role of different modules, the impact of textual instructions, and potential improvements including the use of alternative visual encoders. Our code is available at: https://github.com/yellow-binary-tree/Pfram.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01151
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
Wang, Yueqian
Liang, Jianxin
Wang, Yuxuan
Zhang, Huishuai
Zhao, Dongyan
Computer Vision and Pattern Recognition
Machine Learning
Hallucination is a common issue in Multimodal Large Language Models (MLLMs), yet the underlying principles remain poorly understood. In this paper, we investigate which components of MLLMs contribute to object hallucinations. To analyze image representations while completely avoiding the influence of all other factors other than the image representation itself, we propose a parametric-free representation alignment metric (Pfram) that can measure the similarities between any two representation systems without requiring additional training parameters. Notably, Pfram can also assess the alignment of a neural representation system with the human representation system, represented by ground-truth annotations of images. By evaluating the alignment with object annotations, we demonstrate that this metric shows strong and consistent correlations with object hallucination across a wide range of state-of-the-art MLLMs, spanning various model architectures and sizes. Furthermore, using this metric, we explore other key issues related to image representations in MLLMs, such as the role of different modules, the impact of textual instructions, and potential improvements including the use of alternative visual encoders. Our code is available at: https://github.com/yellow-binary-tree/Pfram.
title Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2409.01151