Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Tianyi, Hu, Zengjie, Sun, Fupeng, Qiu, Jiantao, Jiang, Yizhen, He, Guangxin, Zeng, Bohan, He, Conghui, Yuan, Binhang, Zhang, Wentao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915333163450368
author Bai, Tianyi
Hu, Zengjie
Sun, Fupeng
Qiu, Jiantao
Jiang, Yizhen
He, Guangxin
Zeng, Bohan
He, Conghui
Yuan, Binhang
Zhang, Wentao
author_facet Bai, Tianyi
Hu, Zengjie
Sun, Fupeng
Qiu, Jiantao
Jiang, Yizhen
He, Guangxin
Zeng, Bohan
He, Conghui
Yuan, Binhang
Zhang, Wentao
contents Multi-modal large language models (MLLMs) have achieved remarkable capabilities by integrating visual perception with language understanding, enabling applications such as image-grounded dialogue, visual question answering, and scientific analysis. However, most MLLMs adopt a static inference paradigm, encoding the entire image into fixed visual tokens upfront, which limits their ability to iteratively refine understanding or adapt to context during inference. This contrasts sharply with human perception, which is dynamic, selective, and feedback-driven. In this work, we introduce a novel framework for inference-time visual token scaling that enables MLLMs to perform iterative, verifier-guided reasoning over visual content. We formulate the problem as a Markov Decision Process, involving a reasoner that proposes visual actions and a verifier, which is trained via multi-step Direct Preference Optimization (DPO), that evaluates these actions and determines when reasoning should terminate. To support this, we present a new dataset, VTS, comprising supervised reasoning trajectories (VTS-SFT) and preference-labeled reasoning comparisons (VTS-DPO). Our method significantly outperforms existing approaches across diverse visual reasoning benchmarks, offering not only improved accuracy but also more interpretable and grounded reasoning processes. These results demonstrate the promise of dynamic inference mechanisms for enabling fine-grained, context-aware visual reasoning in next-generation MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
Bai, Tianyi
Hu, Zengjie
Sun, Fupeng
Qiu, Jiantao
Jiang, Yizhen
He, Guangxin
Zeng, Bohan
He, Conghui
Yuan, Binhang
Zhang, Wentao
Computer Vision and Pattern Recognition
Computation and Language
Multi-modal large language models (MLLMs) have achieved remarkable capabilities by integrating visual perception with language understanding, enabling applications such as image-grounded dialogue, visual question answering, and scientific analysis. However, most MLLMs adopt a static inference paradigm, encoding the entire image into fixed visual tokens upfront, which limits their ability to iteratively refine understanding or adapt to context during inference. This contrasts sharply with human perception, which is dynamic, selective, and feedback-driven. In this work, we introduce a novel framework for inference-time visual token scaling that enables MLLMs to perform iterative, verifier-guided reasoning over visual content. We formulate the problem as a Markov Decision Process, involving a reasoner that proposes visual actions and a verifier, which is trained via multi-step Direct Preference Optimization (DPO), that evaluates these actions and determines when reasoning should terminate. To support this, we present a new dataset, VTS, comprising supervised reasoning trajectories (VTS-SFT) and preference-labeled reasoning comparisons (VTS-DPO). Our method significantly outperforms existing approaches across diverse visual reasoning benchmarks, offering not only improved accuracy but also more interpretable and grounded reasoning processes. These results demonstrate the promise of dynamic inference mechanisms for enabling fine-grained, context-aware visual reasoning in next-generation MLLMs.
title Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.07235