When to Think and When to Look: Uncertainty-Guided Lookback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bi, Jing, Bellos, Filippos, Guo, Junjia, Li, Yayuan, Huang, Chao, Tang, Yolo Y., Song, Luchuan, Liang, Susan, Zhang, Zhongfei Mark, Corso, Jason J., Xu, Chenliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911547688747008
author Bi, Jing
Bellos, Filippos
Guo, Junjia
Li, Yayuan
Huang, Chao
Tang, Yolo Y.
Song, Luchuan
Liang, Susan
Zhang, Zhongfei Mark
Corso, Jason J.
Xu, Chenliang
author_facet Bi, Jing
Bellos, Filippos
Guo, Junjia
Li, Yayuan
Huang, Chao
Tang, Yolo Y.
Song, Luchuan
Liang, Susan
Zhang, Zhongfei Mark
Corso, Jason J.
Xu, Chenliang
contents Test-time thinking (that is, generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how thinking actually affects visual reasoning. We provide the first such analysis with a large scale, controlled comparison of thinking for LVLMs, evaluating ten variants from the InternVL3.5 and Qwen3-VL families on MMMU-val under generous token budgets and multi pass decoding. We show that more thinking is not always better; long chains often yield long wrong trajectories that ignore the image and underperform the same models run in standard instruct mode. A deeper analysis reveals that certain short lookback phrases, which explicitly refer back to the image, are strongly enriched in successful trajectories and correlate with better visual grounding. Building on this insight, we propose uncertainty guided lookback, a training free decoding strategy that combines an uncertainty signal with adaptive lookback prompts and breadth search. Our method improves overall MMMU performance, delivers the largest gains in categories where standard thinking is weak, and outperforms several strong decoding baselines, setting a new state of the art under fixed model families and token budgets. We further show that this decoding strategy generalizes, yielding consistent improvements on five additional benchmarks, including two broad multimodal suites and math focused visual reasoning datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When to Think and When to Look: Uncertainty-Guided Lookback
Bi, Jing
Bellos, Filippos
Guo, Junjia
Li, Yayuan
Huang, Chao
Tang, Yolo Y.
Song, Luchuan
Liang, Susan
Zhang, Zhongfei Mark
Corso, Jason J.
Xu, Chenliang
Computer Vision and Pattern Recognition
Computation and Language
Test-time thinking (that is, generating explicit intermediate reasoning chains) is known to boost performance in large language models and has recently shown strong gains for large vision language models (LVLMs). However, despite these promising results, there is still no systematic analysis of how thinking actually affects visual reasoning. We provide the first such analysis with a large scale, controlled comparison of thinking for LVLMs, evaluating ten variants from the InternVL3.5 and Qwen3-VL families on MMMU-val under generous token budgets and multi pass decoding. We show that more thinking is not always better; long chains often yield long wrong trajectories that ignore the image and underperform the same models run in standard instruct mode. A deeper analysis reveals that certain short lookback phrases, which explicitly refer back to the image, are strongly enriched in successful trajectories and correlate with better visual grounding. Building on this insight, we propose uncertainty guided lookback, a training free decoding strategy that combines an uncertainty signal with adaptive lookback prompts and breadth search. Our method improves overall MMMU performance, delivers the largest gains in categories where standard thinking is weak, and outperforms several strong decoding baselines, setting a new state of the art under fixed model families and token budgets. We further show that this decoding strategy generalizes, yielding consistent improvements on five additional benchmarks, including two broad multimodal suites and math focused visual reasoning datasets.
title When to Think and When to Look: Uncertainty-Guided Lookback
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2511.15613