Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tu, Songjun, Zhang, Qichao, Sun, Jingbo, Fu, Yuqian, Li, Linjing, Lan, Xiangyuan, Jiang, Dongmei, Wang, Yaowei, Zhao, Dongbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918148693819392
author Tu, Songjun
Zhang, Qichao
Sun, Jingbo
Fu, Yuqian
Li, Linjing
Lan, Xiangyuan
Jiang, Dongmei
Wang, Yaowei
Zhao, Dongbin
author_facet Tu, Songjun
Zhang, Qichao
Sun, Jingbo
Fu, Yuqian
Li, Linjing
Lan, Xiangyuan
Jiang, Dongmei
Wang, Yaowei
Zhao, Dongbin
contents While multimodal large language models excel at tasks that integrate visual perception with symbolic reasoning, their performance is often undermined by a critical vulnerability: perception-induced errors that propagate through the reasoning chain. Current reinforcement learning (RL) fine-tuning methods, while enhancing reasoning abilities, largely fail to address the underlying misalignment between visual grounding and the subsequent reasoning process. To address this challenge, we propose \textbf{Caption-Regularized Policy Optimization (CapPO)}, a novel RL framework that explicitly enforces perceptual consistency during policy optimization. CapPO integrates two key mechanisms: (1) a caption-based consistency regularization, which minimizes the divergence between responses conditioned on raw images and those conditioned on captions, thereby anchoring reasoning to semantically faithful visual content; and (2) a KL-weighted advantage estimation scheme, which adaptively scales reinforcement signals to strengthen perceptually consistent trajectories while suppressing spurious correlations. Extensive experiments on five math-focused and five general reasoning benchmarks demonstrate that CapPO achieves competitive performance, yielding gains of +6.0% accuracy on math-related tasks and +2.4% on general reasoning tasks over the base Qwen2.5-VL-7B model. Moreover, ablation studies further confirm the effectiveness of each component, while error analysis reveals that CapPO significantly reduces perception-related mistakes compared with baselines. Overall, CapPO provides a simple yet effective framework for improving multimodal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21854
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
Tu, Songjun
Zhang, Qichao
Sun, Jingbo
Fu, Yuqian
Li, Linjing
Lan, Xiangyuan
Jiang, Dongmei
Wang, Yaowei
Zhao, Dongbin
Multimedia
Computer Vision and Pattern Recognition
68T07, 68T45
I.2.6; I.2.7; I.2.10
While multimodal large language models excel at tasks that integrate visual perception with symbolic reasoning, their performance is often undermined by a critical vulnerability: perception-induced errors that propagate through the reasoning chain. Current reinforcement learning (RL) fine-tuning methods, while enhancing reasoning abilities, largely fail to address the underlying misalignment between visual grounding and the subsequent reasoning process. To address this challenge, we propose \textbf{Caption-Regularized Policy Optimization (CapPO)}, a novel RL framework that explicitly enforces perceptual consistency during policy optimization. CapPO integrates two key mechanisms: (1) a caption-based consistency regularization, which minimizes the divergence between responses conditioned on raw images and those conditioned on captions, thereby anchoring reasoning to semantically faithful visual content; and (2) a KL-weighted advantage estimation scheme, which adaptively scales reinforcement signals to strengthen perceptually consistent trajectories while suppressing spurious correlations. Extensive experiments on five math-focused and five general reasoning benchmarks demonstrate that CapPO achieves competitive performance, yielding gains of +6.0% accuracy on math-related tasks and +2.4% on general reasoning tasks over the base Qwen2.5-VL-7B model. Moreover, ablation studies further confirm the effectiveness of each component, while error analysis reveals that CapPO significantly reduces perception-related mistakes compared with baselines. Overall, CapPO provides a simple yet effective framework for improving multimodal reasoning.
title Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
topic Multimedia
Computer Vision and Pattern Recognition
68T07, 68T45
I.2.6; I.2.7; I.2.10
url https://arxiv.org/abs/2509.21854