VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Ming, Wang, Yuanlei, Zhang, Liuzhou, An, Arctanx, Zhang, Renrui, Liang, Hao, Lu, Ming, Shen, Ying, Zhang, Wentao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911282560499712
author Zhong, Ming
Wang, Yuanlei
Zhang, Liuzhou
An, Arctanx
Zhang, Renrui
Liang, Hao
Lu, Ming
Shen, Ying
Zhang, Wentao
author_facet Zhong, Ming
Wang, Yuanlei
Zhang, Liuzhou
An, Arctanx
Zhang, Renrui
Liang, Hao
Lu, Ming
Shen, Ying
Zhang, Wentao
contents While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who naturally bridge details and high-level concepts, models tend to treat these elements in isolation. Prevailing evaluation protocols often decouple low-level perception from high-level reasoning, overlooking their semantic and causal dependencies, which yields non-diagnostic results and obscures performance bottlenecks. We present VCU-Bridge, a framework that operationalizes a human-like hierarchy of visual connotation understanding: multi-level reasoning that advances from foundational perception through semantic bridging to abstract connotation, with an explicit evidence-to-inference trace from concrete cues to abstract conclusions. Building on this framework, we construct HVCU-Bench, a benchmark for hierarchical visual connotation understanding with explicit, level-wise diagnostics. Comprehensive experiments demonstrate a consistent decline in performance as reasoning progresses to higher levels. We further develop a data generation pipeline for instruction tuning guided by Monte Carlo Tree Search (MCTS) and show that strengthening low-level capabilities yields measurable gains at higher levels. Interestingly, it not only improves on HVCU-Bench but also brings benefits on general benchmarks (average +2.53%), especially with substantial gains on MMStar (+7.26%), demonstrating the significance of the hierarchical thinking pattern and its effectiveness in enhancing MLLM capabilities. The project page is at https://vcu-bridge.github.io .
format Preprint
id arxiv_https___arxiv_org_abs_2511_18121
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
Zhong, Ming
Wang, Yuanlei
Zhang, Liuzhou
An, Arctanx
Zhang, Renrui
Liang, Hao
Lu, Ming
Shen, Ying
Zhang, Wentao
Computer Vision and Pattern Recognition
Artificial Intelligence
While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who naturally bridge details and high-level concepts, models tend to treat these elements in isolation. Prevailing evaluation protocols often decouple low-level perception from high-level reasoning, overlooking their semantic and causal dependencies, which yields non-diagnostic results and obscures performance bottlenecks. We present VCU-Bridge, a framework that operationalizes a human-like hierarchy of visual connotation understanding: multi-level reasoning that advances from foundational perception through semantic bridging to abstract connotation, with an explicit evidence-to-inference trace from concrete cues to abstract conclusions. Building on this framework, we construct HVCU-Bench, a benchmark for hierarchical visual connotation understanding with explicit, level-wise diagnostics. Comprehensive experiments demonstrate a consistent decline in performance as reasoning progresses to higher levels. We further develop a data generation pipeline for instruction tuning guided by Monte Carlo Tree Search (MCTS) and show that strengthening low-level capabilities yields measurable gains at higher levels. Interestingly, it not only improves on HVCU-Bench but also brings benefits on general benchmarks (average +2.53%), especially with substantial gains on MMStar (+7.26%), demonstrating the significance of the hierarchical thinking pattern and its effectiveness in enhancing MLLM capabilities. The project page is at https://vcu-bridge.github.io .
title VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.18121