LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909499446525952 |
|---|---|
| author | Wu, Zongyu Niu, Yuwei Gao, Hongcheng Lin, Minhua Zhang, Zhiwei Zhang, Zhifang Shi, Qi Wang, Yilong Fu, Sike Xu, Junjie Ao, Junjie Dai, Enyan Feng, Lei Zhang, Xiang Wang, Suhang |
| author_facet | Wu, Zongyu Niu, Yuwei Gao, Hongcheng Lin, Minhua Zhang, Zhiwei Zhang, Zhifang Shi, Qi Wang, Yilong Fu, Sike Xu, Junjie Ao, Junjie Dai, Enyan Feng, Lei Zhang, Xiang Wang, Suhang |
| contents | Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders their adoption in the real world. Existing studies emphasized that the strong language priors of LVLMs can overpower visual information, causing hallucinations. However, the positive role of language priors is the key to a powerful LVLM. If the language priors are too weak, LVLMs will struggle to leverage rich parameter knowledge and instruction understanding abilities to complete tasks in challenging visual scenarios where visual information alone is insufficient. Therefore, we propose a benchmark called LanP to rethink the impact of Language Priors in LVLMs. It is designed to investigate how strong language priors are in current LVLMs. LanP consists of 170 images and 340 corresponding well-designed questions. Extensive experiments on 25 popular LVLMs reveal that many LVLMs' language priors are not strong enough to effectively aid question answering when objects are partially hidden. Many models, including GPT-4 Turbo, exhibit an accuracy below 0.5 in such a scenario. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_12359 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models Wu, Zongyu Niu, Yuwei Gao, Hongcheng Lin, Minhua Zhang, Zhiwei Zhang, Zhifang Shi, Qi Wang, Yilong Fu, Sike Xu, Junjie Ao, Junjie Dai, Enyan Feng, Lei Zhang, Xiang Wang, Suhang Computer Vision and Pattern Recognition Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders their adoption in the real world. Existing studies emphasized that the strong language priors of LVLMs can overpower visual information, causing hallucinations. However, the positive role of language priors is the key to a powerful LVLM. If the language priors are too weak, LVLMs will struggle to leverage rich parameter knowledge and instruction understanding abilities to complete tasks in challenging visual scenarios where visual information alone is insufficient. Therefore, we propose a benchmark called LanP to rethink the impact of Language Priors in LVLMs. It is designed to investigate how strong language priors are in current LVLMs. LanP consists of 170 images and 340 corresponding well-designed questions. Extensive experiments on 25 popular LVLMs reveal that many LVLMs' language priors are not strong enough to effectively aid question answering when objects are partially hidden. Many models, including GPT-4 Turbo, exhibit an accuracy below 0.5 in such a scenario. |
| title | LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2502.12359 |