Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Siqi, Gao, Zilve, Qiu, Haibo, Liu, Fanfan, Shi, Peng, Zeng, Zhixiong, Liao, Qingmin, Ma, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917156862558208
author Yang, Siqi
Gao, Zilve
Qiu, Haibo
Liu, Fanfan
Shi, Peng
Zeng, Zhixiong
Liao, Qingmin
Ma, Lin
author_facet Yang, Siqi
Gao, Zilve
Qiu, Haibo
Liu, Fanfan
Shi, Peng
Zeng, Zhixiong
Liao, Qingmin
Ma, Lin
contents Multimodal Large Language Models (MLLMs) demonstrate significant potential but remain brittle in complex, long-chain visual reasoning tasks. A critical failure mode is "visual forgetting", where models progressively lose visual grounding as reasoning extends, a phenomenon aptly described as "think longer, see less". We posit this failure stems from current training paradigms prematurely entangling two distinct cognitive skills: (1) abstract logical reasoning "how-to-think") and (2) strategic visual perception ("when-to-look"). This creates a foundational cold-start deficiency -- weakening abstract reasoning -- and a strategic perception deficit, as models lack a policy for when to perceive. In this paper, we propose a novel curriculum-based framework to disentangle these skills. First, we introduce a disentangled Supervised Fine-Tuning (SFT) curriculum that builds a robust abstract reasoning backbone on text-only data before anchoring it to vision with a novel Perception-Grounded Chain-of-Thought (PG-CoT) paradigm. Second, we resolve the strategic perception deficit by formulating timing as a reinforcement learning problem. We design a Pivotal Perception Reward that teaches the model when to look by coupling perceptual actions to linguistic markers of cognitive uncertainty (e.g., "wait", "verify"), thereby learning an autonomous grounding policy. Our contributions include the formalization of these two deficiencies and the development of a principled, two-stage framework to address them, transforming the model from a heuristic-driven observer to a strategic, grounded reasoner. \textbf{Code}: \url{https://github.com/gaozilve-max/learning-when-to-look}.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
Yang, Siqi
Gao, Zilve
Qiu, Haibo
Liu, Fanfan
Shi, Peng
Zeng, Zhixiong
Liao, Qingmin
Ma, Lin
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) demonstrate significant potential but remain brittle in complex, long-chain visual reasoning tasks. A critical failure mode is "visual forgetting", where models progressively lose visual grounding as reasoning extends, a phenomenon aptly described as "think longer, see less". We posit this failure stems from current training paradigms prematurely entangling two distinct cognitive skills: (1) abstract logical reasoning "how-to-think") and (2) strategic visual perception ("when-to-look"). This creates a foundational cold-start deficiency -- weakening abstract reasoning -- and a strategic perception deficit, as models lack a policy for when to perceive. In this paper, we propose a novel curriculum-based framework to disentangle these skills. First, we introduce a disentangled Supervised Fine-Tuning (SFT) curriculum that builds a robust abstract reasoning backbone on text-only data before anchoring it to vision with a novel Perception-Grounded Chain-of-Thought (PG-CoT) paradigm. Second, we resolve the strategic perception deficit by formulating timing as a reinforcement learning problem. We design a Pivotal Perception Reward that teaches the model when to look by coupling perceptual actions to linguistic markers of cognitive uncertainty (e.g., "wait", "verify"), thereby learning an autonomous grounding policy. Our contributions include the formalization of these two deficiencies and the development of a principled, two-stage framework to address them, transforming the model from a heuristic-driven observer to a strategic, grounded reasoner. \textbf{Code}: \url{https://github.com/gaozilve-max/learning-when-to-look}.
title Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17227