From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Ruilin, Shi, Chufan, Zhang, Yizhen, Yang, Cheng, Jiang, Songtao, Guan, Tongkun, Chen, Ruizhe, Chu, Ruihang, Wang, Peng, Yang, Mingkun, Yang, Yujiu, Lin, Junyang, Yang, Zhibo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915833044795392
author Luo, Ruilin
Shi, Chufan
Zhang, Yizhen
Yang, Cheng
Jiang, Songtao
Guan, Tongkun
Chen, Ruizhe
Chu, Ruihang
Wang, Peng
Yang, Mingkun
Yang, Yujiu
Lin, Junyang
Yang, Zhibo
author_facet Luo, Ruilin
Shi, Chufan
Zhang, Yizhen
Yang, Cheng
Jiang, Songtao
Guan, Tongkun
Chen, Ruizhe
Chu, Ruihang
Wang, Peng
Yang, Mingkun
Yang, Yujiu
Lin, Junyang
Yang, Zhibo
contents The cold-start initialization stage plays a pivotal role in training Multimodal Large Reasoning Models (MLRMs), yet its mechanisms remain insufficiently understood. To analyze this stage, we introduce the Visual Attention Score (VAS), an attention-based metric that quantifies how much a model attends to visual tokens. We find that reasoning performance is strongly correlated with VAS (r=0.9616): models with higher VAS achieve substantially stronger multimodal reasoning. Surprisingly, multimodal cold-start fails to elevate VAS, resulting in attention distributions close to the base model, whereas text-only cold-start leads to a clear increase. We term this counter-intuitive phenomenon Lazy Attention Localization. To validate its causal role, we design training-free interventions that directly modulate attention allocation during inference, performance gains of 1$-$2% without any retraining. Building on these insights, we further propose Attention-Guided Visual Anchoring and Reflection (AVAR), a comprehensive cold-start framework that integrates visual-anchored data synthesis, attention-guided objectives, and visual-anchored reward shaping. Applied to Qwen2.5-VL-7B, AVAR achieves an average gain of 7.0% across 7 multimodal reasoning benchmarks. Ablation studies further confirm that each component of AVAR contributes step-wise to the overall gains. The code, data, and models are available at https://github.com/lrlbbzl/Qwen-AVAR.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03825
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
Luo, Ruilin
Shi, Chufan
Zhang, Yizhen
Yang, Cheng
Jiang, Songtao
Guan, Tongkun
Chen, Ruizhe
Chu, Ruihang
Wang, Peng
Yang, Mingkun
Yang, Yujiu
Lin, Junyang
Yang, Zhibo
Computer Vision and Pattern Recognition
Artificial Intelligence
The cold-start initialization stage plays a pivotal role in training Multimodal Large Reasoning Models (MLRMs), yet its mechanisms remain insufficiently understood. To analyze this stage, we introduce the Visual Attention Score (VAS), an attention-based metric that quantifies how much a model attends to visual tokens. We find that reasoning performance is strongly correlated with VAS (r=0.9616): models with higher VAS achieve substantially stronger multimodal reasoning. Surprisingly, multimodal cold-start fails to elevate VAS, resulting in attention distributions close to the base model, whereas text-only cold-start leads to a clear increase. We term this counter-intuitive phenomenon Lazy Attention Localization. To validate its causal role, we design training-free interventions that directly modulate attention allocation during inference, performance gains of 1$-$2% without any retraining. Building on these insights, we further propose Attention-Guided Visual Anchoring and Reflection (AVAR), a comprehensive cold-start framework that integrates visual-anchored data synthesis, attention-guided objectives, and visual-anchored reward shaping. Applied to Qwen2.5-VL-7B, AVAR achieves an average gain of 7.0% across 7 multimodal reasoning benchmarks. Ablation studies further confirm that each component of AVAR contributes step-wise to the overall gains. The code, data, and models are available at https://github.com/lrlbbzl/Qwen-AVAR.
title From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.03825