Where do Large Vision-Language Models Look at when Answering Questions?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xing, Xiaoying, Kuo, Chia-Wen, Fuxin, Li, Niu, Yulei, Chen, Fan, Li, Ming, Wu, Ying, Wen, Longyin, Zhu, Sijie
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910881753858048
author Xing, Xiaoying
Kuo, Chia-Wen
Fuxin, Li
Niu, Yulei
Chen, Fan
Li, Ming
Wu, Ying
Wen, Longyin
Zhu, Sijie
author_facet Xing, Xiaoying
Kuo, Chia-Wen
Fuxin, Li
Niu, Yulei
Chen, Fan
Li, Ming
Wu, Ying
Wen, Longyin
Zhu, Sijie
contents Large Vision-Language Models (LVLMs) have shown promising performance in vision-language understanding and reasoning tasks. However, their visual understanding behaviors remain underexplored. A fundamental question arises: to what extent do LVLMs rely on visual input, and which image regions contribute to their responses? It is non-trivial to interpret the free-form generation of LVLMs due to their complicated visual architecture (e.g., multiple encoders and multi-resolution) and variable-length outputs. In this paper, we extend existing heatmap visualization methods (e.g., iGOS++) to support LVLMs for open-ended visual question answering. We propose a method to select visually relevant tokens that reflect the relevance between generated answers and input image. Furthermore, we conduct a comprehensive analysis of state-of-the-art LVLMs on benchmarks designed to require visual information to answer. Our findings offer several insights into LVLM behavior, including the relationship between focus region and answer correctness, differences in visual attention across architectures, and the impact of LLM scale on visual understanding. The code and data are available at https://github.com/bytedance/LVLM_Interpretation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Where do Large Vision-Language Models Look at when Answering Questions?
Xing, Xiaoying
Kuo, Chia-Wen
Fuxin, Li
Niu, Yulei
Chen, Fan
Li, Ming
Wu, Ying
Wen, Longyin
Zhu, Sijie
Computer Vision and Pattern Recognition
Computation and Language
Large Vision-Language Models (LVLMs) have shown promising performance in vision-language understanding and reasoning tasks. However, their visual understanding behaviors remain underexplored. A fundamental question arises: to what extent do LVLMs rely on visual input, and which image regions contribute to their responses? It is non-trivial to interpret the free-form generation of LVLMs due to their complicated visual architecture (e.g., multiple encoders and multi-resolution) and variable-length outputs. In this paper, we extend existing heatmap visualization methods (e.g., iGOS++) to support LVLMs for open-ended visual question answering. We propose a method to select visually relevant tokens that reflect the relevance between generated answers and input image. Furthermore, we conduct a comprehensive analysis of state-of-the-art LVLMs on benchmarks designed to require visual information to answer. Our findings offer several insights into LVLM behavior, including the relationship between focus region and answer correctness, differences in visual attention across architectures, and the impact of LLM scale on visual understanding. The code and data are available at https://github.com/bytedance/LVLM_Interpretation.
title Where do Large Vision-Language Models Look at when Answering Questions?
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.13891