VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Yizhuo, Chen, Mingkang, Feng, Zhibang, Xiao, Tong, Qu, Wanying, Shao, Wenqi, Fu, Yanwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912615068860416
author Ding, Yizhuo
Chen, Mingkang
Feng, Zhibang
Xiao, Tong
Qu, Wanying
Shao, Wenqi
Fu, Yanwei
author_facet Ding, Yizhuo
Chen, Mingkang
Feng, Zhibang
Xiao, Tong
Qu, Wanying
Shao, Wenqi
Fu, Yanwei
contents Multimodal large language models (MLLMs) often struggle to ground reasoning in perceptual evidence. We present a systematic study of perception strategies-explicit, implicit, visual, and textual-across four multimodal benchmarks and two MLLMs. Our findings show that explicit perception, especially when paired with textual cues, consistently yields the best improvements, particularly for smaller models. Based on this insight, we propose VTPerception-R1, a unified two-stage framework that decouples perception from reasoning. Stage 1 introduces perception-augmented fine-tuning, and Stage 2 applies perception-aware reinforcement learning with novel visual, textual, and consistency rewards. Experiments demonstrate that VTPerception-R1 significantly improves reasoning accuracy and robustness across diverse tasks, offering a scalable and auditable solution for perception-grounded multimodal reasoning. Our code is available at: https://github.com/yizhuoDi/VTPerceprion-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
Ding, Yizhuo
Chen, Mingkang
Feng, Zhibang
Xiao, Tong
Qu, Wanying
Shao, Wenqi
Fu, Yanwei
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal large language models (MLLMs) often struggle to ground reasoning in perceptual evidence. We present a systematic study of perception strategies-explicit, implicit, visual, and textual-across four multimodal benchmarks and two MLLMs. Our findings show that explicit perception, especially when paired with textual cues, consistently yields the best improvements, particularly for smaller models. Based on this insight, we propose VTPerception-R1, a unified two-stage framework that decouples perception from reasoning. Stage 1 introduces perception-augmented fine-tuning, and Stage 2 applies perception-aware reinforcement learning with novel visual, textual, and consistency rewards. Experiments demonstrate that VTPerception-R1 significantly improves reasoning accuracy and robustness across diverse tasks, offering a scalable and auditable solution for perception-grounded multimodal reasoning. Our code is available at: https://github.com/yizhuoDi/VTPerceprion-R1.
title VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.24776