LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zongyu, Niu, Yuwei, Gao, Hongcheng, Lin, Minhua, Zhang, Zhiwei, Zhang, Zhifang, Shi, Qi, Wang, Yilong, Fu, Sike, Xu, Junjie, Ao, Junjie, Dai, Enyan, Feng, Lei, Zhang, Xiang, Wang, Suhang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909499446525952
author Wu, Zongyu
Niu, Yuwei
Gao, Hongcheng
Lin, Minhua
Zhang, Zhiwei
Zhang, Zhifang
Shi, Qi
Wang, Yilong
Fu, Sike
Xu, Junjie
Ao, Junjie
Dai, Enyan
Feng, Lei
Zhang, Xiang
Wang, Suhang
author_facet Wu, Zongyu
Niu, Yuwei
Gao, Hongcheng
Lin, Minhua
Zhang, Zhiwei
Zhang, Zhifang
Shi, Qi
Wang, Yilong
Fu, Sike
Xu, Junjie
Ao, Junjie
Dai, Enyan
Feng, Lei
Zhang, Xiang
Wang, Suhang
contents Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders their adoption in the real world. Existing studies emphasized that the strong language priors of LVLMs can overpower visual information, causing hallucinations. However, the positive role of language priors is the key to a powerful LVLM. If the language priors are too weak, LVLMs will struggle to leverage rich parameter knowledge and instruction understanding abilities to complete tasks in challenging visual scenarios where visual information alone is insufficient. Therefore, we propose a benchmark called LanP to rethink the impact of Language Priors in LVLMs. It is designed to investigate how strong language priors are in current LVLMs. LanP consists of 170 images and 340 corresponding well-designed questions. Extensive experiments on 25 popular LVLMs reveal that many LVLMs' language priors are not strong enough to effectively aid question answering when objects are partially hidden. Many models, including GPT-4 Turbo, exhibit an accuracy below 0.5 in such a scenario.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12359
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
Wu, Zongyu
Niu, Yuwei
Gao, Hongcheng
Lin, Minhua
Zhang, Zhiwei
Zhang, Zhifang
Shi, Qi
Wang, Yilong
Fu, Sike
Xu, Junjie
Ao, Junjie
Dai, Enyan
Feng, Lei
Zhang, Xiang
Wang, Suhang
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders their adoption in the real world. Existing studies emphasized that the strong language priors of LVLMs can overpower visual information, causing hallucinations. However, the positive role of language priors is the key to a powerful LVLM. If the language priors are too weak, LVLMs will struggle to leverage rich parameter knowledge and instruction understanding abilities to complete tasks in challenging visual scenarios where visual information alone is insufficient. Therefore, we propose a benchmark called LanP to rethink the impact of Language Priors in LVLMs. It is designed to investigate how strong language priors are in current LVLMs. LanP consists of 170 images and 340 corresponding well-designed questions. Extensive experiments on 25 popular LVLMs reveal that many LVLMs' language priors are not strong enough to effectively aid question answering when objects are partially hidden. Many models, including GPT-4 Turbo, exhibit an accuracy below 0.5 in such a scenario.
title LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.12359