ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: You, Weihang, Zhu, Qingchan, Liu, David, Pan, Yi, Yuan, Geng, Jiang, Hanqi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914303534170112
author You, Weihang
Zhu, Qingchan
Liu, David
Pan, Yi
Yuan, Geng
Jiang, Hanqi
author_facet You, Weihang
Zhu, Qingchan
Liu, David
Pan, Yi
Yuan, Geng
Jiang, Hanqi
contents Chain-of-Thought (CoT) reasoning excels in language models but struggles in vision-language models due to premature visual-to-text conversion that discards continuous information such as geometry and spatial layout. While recent methods enhance CoT through static enumeration or attention-based selection, they remain passive, i.e., processing pre-computed inputs rather than actively seeking task-relevant details. Inspired by human active perception, we introduce ViThinker, a framework that enables vision-language models to autonomously generate decision (query) tokens triggering the synthesis of expert-aligned visual features on demand. ViThinker internalizes vision-expert capabilities during training, performing generative mental simulation during inference without external tool calls. Through a two-stage curriculum: first distilling frozen experts into model parameters, then learning task-driven querying via sparsity penalties, i.e., ViThinker discovers minimal sufficient perception for each reasoning step. Evaluations across vision-centric benchmarks demonstrate consistent improvements, validating that active query generation outperforms passive approaches in both perceptual grounding and reasoning accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02873
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying
You, Weihang
Zhu, Qingchan
Liu, David
Pan, Yi
Yuan, Geng
Jiang, Hanqi
Computer Vision and Pattern Recognition
Chain-of-Thought (CoT) reasoning excels in language models but struggles in vision-language models due to premature visual-to-text conversion that discards continuous information such as geometry and spatial layout. While recent methods enhance CoT through static enumeration or attention-based selection, they remain passive, i.e., processing pre-computed inputs rather than actively seeking task-relevant details. Inspired by human active perception, we introduce ViThinker, a framework that enables vision-language models to autonomously generate decision (query) tokens triggering the synthesis of expert-aligned visual features on demand. ViThinker internalizes vision-expert capabilities during training, performing generative mental simulation during inference without external tool calls. Through a two-stage curriculum: first distilling frozen experts into model parameters, then learning task-driven querying via sparsity penalties, i.e., ViThinker discovers minimal sufficient perception for each reasoning step. Evaluations across vision-centric benchmarks demonstrate consistent improvements, validating that active query generation outperforms passive approaches in both perceptual grounding and reasoning accuracy.
title ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.02873