PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Cong, Liu, Mingyu, Jing, Chenchen, Zhou, Yizhou, Rao, Fengyun, Chen, Hao, Zhang, Bo, Shen, Chunhua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929748924432384
author Chen, Cong
Liu, Mingyu
Jing, Chenchen
Zhou, Yizhou
Rao, Fengyun
Chen, Hao
Zhang, Bo
Shen, Chunhua
author_facet Chen, Cong
Liu, Mingyu
Jing, Chenchen
Zhou, Yizhou
Rao, Fengyun
Chen, Hao
Zhang, Bo
Shen, Chunhua
contents This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalFscore, a novel metric built upon the language graph and is designed to evaluate both the accuracy and completeness of dense captions at a granular level. Additionally, we identify the root cause of hallucination as the model's over-reliance on its language prior. To address this, we propose PerturboLLaVA, which reduces the model's reliance on the language prior by incorporating adversarially perturbed text during training. This method enhances the model's focus on visual inputs, effectively reducing hallucinations and producing accurate, image-grounded descriptions without incurring additional computational overhead. PerturboLLaVA significantly improves the fidelity of generated captions, outperforming existing approaches in handling multimodal hallucinations and achieving improved performance across general multimodal benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06486
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
Chen, Cong
Liu, Mingyu
Jing, Chenchen
Zhou, Yizhou
Rao, Fengyun
Chen, Hao
Zhang, Bo
Shen, Chunhua
Computer Vision and Pattern Recognition
Artificial Intelligence
This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalFscore, a novel metric built upon the language graph and is designed to evaluate both the accuracy and completeness of dense captions at a granular level. Additionally, we identify the root cause of hallucination as the model's over-reliance on its language prior. To address this, we propose PerturboLLaVA, which reduces the model's reliance on the language prior by incorporating adversarially perturbed text during training. This method enhances the model's focus on visual inputs, effectively reducing hallucinations and producing accurate, image-grounded descriptions without incurring additional computational overhead. PerturboLLaVA significantly improves the fidelity of generated captions, outperforming existing approaches in handling multimodal hallucinations and achieving improved performance across general multimodal benchmarks.
title PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.06486