Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Pengfei, Xu, Guohai, Wang, Weinong, Yang, Junjie, Lou, Jie, Xue, Yunhua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2505.10541
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916753976590336
author Wang, Pengfei
Xu, Guohai
Wang, Weinong
Yang, Junjie
Lou, Jie
Xue, Yunhua
author_facet Wang, Pengfei
Xu, Guohai
Wang, Weinong
Yang, Junjie
Lou, Jie
Xue, Yunhua
contents Recent advancements have enhanced the capability of Multimodal Large Language Models (MLLMs) to comprehend multi-image information. However, existing benchmarks primarily evaluate answer correctness, overlooking whether models genuinely comprehend the visual input. To address this, we define implicit visual misunderstanding (IVM), where MLLMs provide correct answers without fully comprehending the visual input. Through our analysis, we decouple the visual and textual modalities within the causal attention module, revealing that attention distribution increasingly converges on the image associated with the correct answer as the network layers deepen. This insight leads to the introduction of a scale-agnostic metric, \textit{attention accuracy}, and a novel benchmark for quantifying IVMs. Attention accuracy directly evaluates the model's visual understanding via internal mechanisms, remaining robust to positional biases for more reliable assessments. Furthermore, we extend our approach to finer granularities and demonstrate its effectiveness in unimodal scenarios, underscoring its versatility and generalizability.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10541
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
Wang, Pengfei
Xu, Guohai
Wang, Weinong
Yang, Junjie
Lou, Jie
Xue, Yunhua
Computer Vision and Pattern Recognition
Recent advancements have enhanced the capability of Multimodal Large Language Models (MLLMs) to comprehend multi-image information. However, existing benchmarks primarily evaluate answer correctness, overlooking whether models genuinely comprehend the visual input. To address this, we define implicit visual misunderstanding (IVM), where MLLMs provide correct answers without fully comprehending the visual input. Through our analysis, we decouple the visual and textual modalities within the causal attention module, revealing that attention distribution increasingly converges on the image associated with the correct answer as the network layers deepen. This insight leads to the introduction of a scale-agnostic metric, \textit{attention accuracy}, and a novel benchmark for quantifying IVMs. Attention accuracy directly evaluates the model's visual understanding via internal mechanisms, remaining robust to positional biases for more reliable assessments. Furthermore, we extend our approach to finer granularities and demonstrate its effectiveness in unimodal scenarios, underscoring its versatility and generalizability.
title Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.10541