Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918148693819392 |
|---|---|
| author | Tu, Songjun Zhang, Qichao Sun, Jingbo Fu, Yuqian Li, Linjing Lan, Xiangyuan Jiang, Dongmei Wang, Yaowei Zhao, Dongbin |
| author_facet | Tu, Songjun Zhang, Qichao Sun, Jingbo Fu, Yuqian Li, Linjing Lan, Xiangyuan Jiang, Dongmei Wang, Yaowei Zhao, Dongbin |
| contents | While multimodal large language models excel at tasks that integrate visual perception with symbolic reasoning, their performance is often undermined by a critical vulnerability: perception-induced errors that propagate through the reasoning chain. Current reinforcement learning (RL) fine-tuning methods, while enhancing reasoning abilities, largely fail to address the underlying misalignment between visual grounding and the subsequent reasoning process. To address this challenge, we propose \textbf{Caption-Regularized Policy Optimization (CapPO)}, a novel RL framework that explicitly enforces perceptual consistency during policy optimization. CapPO integrates two key mechanisms: (1) a caption-based consistency regularization, which minimizes the divergence between responses conditioned on raw images and those conditioned on captions, thereby anchoring reasoning to semantically faithful visual content; and (2) a KL-weighted advantage estimation scheme, which adaptively scales reinforcement signals to strengthen perceptually consistent trajectories while suppressing spurious correlations. Extensive experiments on five math-focused and five general reasoning benchmarks demonstrate that CapPO achieves competitive performance, yielding gains of +6.0% accuracy on math-related tasks and +2.4% on general reasoning tasks over the base Qwen2.5-VL-7B model. Moreover, ablation studies further confirm the effectiveness of each component, while error analysis reveals that CapPO significantly reduces perception-related mistakes compared with baselines. Overall, CapPO provides a simple yet effective framework for improving multimodal reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_21854 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization Tu, Songjun Zhang, Qichao Sun, Jingbo Fu, Yuqian Li, Linjing Lan, Xiangyuan Jiang, Dongmei Wang, Yaowei Zhao, Dongbin Multimedia Computer Vision and Pattern Recognition 68T07, 68T45 I.2.6; I.2.7; I.2.10 While multimodal large language models excel at tasks that integrate visual perception with symbolic reasoning, their performance is often undermined by a critical vulnerability: perception-induced errors that propagate through the reasoning chain. Current reinforcement learning (RL) fine-tuning methods, while enhancing reasoning abilities, largely fail to address the underlying misalignment between visual grounding and the subsequent reasoning process. To address this challenge, we propose \textbf{Caption-Regularized Policy Optimization (CapPO)}, a novel RL framework that explicitly enforces perceptual consistency during policy optimization. CapPO integrates two key mechanisms: (1) a caption-based consistency regularization, which minimizes the divergence between responses conditioned on raw images and those conditioned on captions, thereby anchoring reasoning to semantically faithful visual content; and (2) a KL-weighted advantage estimation scheme, which adaptively scales reinforcement signals to strengthen perceptually consistent trajectories while suppressing spurious correlations. Extensive experiments on five math-focused and five general reasoning benchmarks demonstrate that CapPO achieves competitive performance, yielding gains of +6.0% accuracy on math-related tasks and +2.4% on general reasoning tasks over the base Qwen2.5-VL-7B model. Moreover, ablation studies further confirm the effectiveness of each component, while error analysis reveals that CapPO significantly reduces perception-related mistakes compared with baselines. Overall, CapPO provides a simple yet effective framework for improving multimodal reasoning. |
| title | Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization |
| topic | Multimedia Computer Vision and Pattern Recognition 68T07, 68T45 I.2.6; I.2.7; I.2.10 |
| url | https://arxiv.org/abs/2509.21854 |