Hallucination Augmented Contrastive Learning for Multimodal Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Chaoya, Xu, Haiyang, Dong, Mengfan, Chen, Jiaxing, Ye, Wei, Yan, Ming, Ye, Qinghao, Zhang, Ji, Huang, Fei, Zhang, Shikun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913243304296448
author Jiang, Chaoya
Xu, Haiyang
Dong, Mengfan
Chen, Jiaxing
Ye, Wei
Yan, Ming
Ye, Qinghao
Zhang, Ji
Huang, Fei
Zhang, Shikun
author_facet Jiang, Chaoya
Xu, Haiyang
Dong, Mengfan
Chen, Jiaxing
Ye, Wei
Yan, Ming
Ye, Qinghao
Zhang, Ji
Huang, Fei
Zhang, Shikun
contents Multi-modal large language models (MLLMs) have been shown to efficiently integrate natural language with visual information to handle multi-modal tasks. However, MLLMs still face a fundamental limitation of hallucinations, where they tend to generate erroneous or fabricated information. In this paper, we address hallucinations in MLLMs from a novel perspective of representation learning. We first analyzed the representation distribution of textual and visual tokens in MLLM, revealing two important findings: 1) there is a significant gap between textual and visual representations, indicating unsatisfactory cross-modal representation alignment; 2) representations of texts that contain and do not contain hallucinations are entangled, making it challenging to distinguish them. These two observations inspire us with a simple yet effective method to mitigate hallucinations. Specifically, we introduce contrastive learning into MLLMs and use text with hallucination as hard negative examples, naturally bringing representations of non-hallucinative text and visual samples closer while pushing way representations of non-hallucinating and hallucinative text. We evaluate our method quantitatively and qualitatively, showing its effectiveness in reducing hallucination occurrences and improving performance across multiple benchmarks. On the MMhal-Bench benchmark, our method obtains a 34.66% /29.5% improvement over the baseline MiniGPT-4/LLaVA. Our code is available on https://github.com/X-PLUG/mPLUG-HalOwl/tree/main/hacl.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06968
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
Jiang, Chaoya
Xu, Haiyang
Dong, Mengfan
Chen, Jiaxing
Ye, Wei
Yan, Ming
Ye, Qinghao
Zhang, Ji
Huang, Fei
Zhang, Shikun
Computer Vision and Pattern Recognition
Multi-modal large language models (MLLMs) have been shown to efficiently integrate natural language with visual information to handle multi-modal tasks. However, MLLMs still face a fundamental limitation of hallucinations, where they tend to generate erroneous or fabricated information. In this paper, we address hallucinations in MLLMs from a novel perspective of representation learning. We first analyzed the representation distribution of textual and visual tokens in MLLM, revealing two important findings: 1) there is a significant gap between textual and visual representations, indicating unsatisfactory cross-modal representation alignment; 2) representations of texts that contain and do not contain hallucinations are entangled, making it challenging to distinguish them. These two observations inspire us with a simple yet effective method to mitigate hallucinations. Specifically, we introduce contrastive learning into MLLMs and use text with hallucination as hard negative examples, naturally bringing representations of non-hallucinative text and visual samples closer while pushing way representations of non-hallucinating and hallucinative text. We evaluate our method quantitatively and qualitatively, showing its effectiveness in reducing hallucination occurrences and improving performance across multiple benchmarks. On the MMhal-Bench benchmark, our method obtains a 34.66% /29.5% improvement over the baseline MiniGPT-4/LLaVA. Our code is available on https://github.com/X-PLUG/mPLUG-HalOwl/tree/main/hacl.
title Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.06968