More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tian, Xinyu, Zou, Shu, Yang, Zhaoyuan, He, Mengqi, Waschkowski, Fabian, Wesemann, Lukas, Tu, Peter, Zhang, Jing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908918334095360
author Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
He, Mengqi
Waschkowski, Fabian
Wesemann, Lukas
Tu, Peter
Zhang, Jing
author_facet Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
He, Mengqi
Waschkowski, Fabian
Wesemann, Lukas
Tu, Peter
Zhang, Jing
contents Reasoning has emerged as a pivotal capability in Large Language Models (LLMs). Through Reinforcement Learning (RL), typically Group Relative Policy Optimization (GRPO), these models are able to solve complex tasks such as mathematics and code generation. Building on these advances, recent research has sought to extend reasoning to Vision-Language Models (VLMs), yielding promising results across diverse visual tasks. Despite this progress, our study uncovers the dual nature of multimodal reasoning: while it substantially enhances logical inference and facilitates performance on challenging problems, it may gradually impair perceptual grounding, leading to recognition failures on otherwise basic visual questions. Through further analysis, we attribute this phenomenon to visual forgetting, wherein prolonged reasoning causes the model to increasingly disregard visual input. To address this, we propose Vision-Anchored Policy Optimization (VAPO), a simple yet effective method that explicitly steers the reasoning process toward visually grounded trajectories. Our result model, VAPO-Thinker-7B, significantly strengthens the model's reliance on visual information and achieves new state-of-the-art results on a wide range of established benchmarks. Project page: https://xytian1008.github.io/VAPO/
format Preprint
id arxiv_https___arxiv_org_abs_2509_25848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
He, Mengqi
Waschkowski, Fabian
Wesemann, Lukas
Tu, Peter
Zhang, Jing
Computer Vision and Pattern Recognition
Artificial Intelligence
Reasoning has emerged as a pivotal capability in Large Language Models (LLMs). Through Reinforcement Learning (RL), typically Group Relative Policy Optimization (GRPO), these models are able to solve complex tasks such as mathematics and code generation. Building on these advances, recent research has sought to extend reasoning to Vision-Language Models (VLMs), yielding promising results across diverse visual tasks. Despite this progress, our study uncovers the dual nature of multimodal reasoning: while it substantially enhances logical inference and facilitates performance on challenging problems, it may gradually impair perceptual grounding, leading to recognition failures on otherwise basic visual questions. Through further analysis, we attribute this phenomenon to visual forgetting, wherein prolonged reasoning causes the model to increasingly disregard visual input. To address this, we propose Vision-Anchored Policy Optimization (VAPO), a simple yet effective method that explicitly steers the reasoning process toward visually grounded trajectories. Our result model, VAPO-Thinker-7B, significantly strengthens the model's reliance on visual information and achieves new state-of-the-art results on a wide range of established benchmarks. Project page: https://xytian1008.github.io/VAPO/
title More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.25848