Stable and Efficient Single-Rollout RL for Multimodal Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Rui, Yu, Dian, Ke, Lei, Liu, Haolin, Zhou, Yujun, Liang, Zhenwen, Mi, Haitao, Tokekar, Pratap, Yu, Dong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918257008574464
author Liu, Rui
Yu, Dian
Ke, Lei
Liu, Haolin
Zhou, Yujun
Liang, Zhenwen
Mi, Haitao
Tokekar, Pratap
Yu, Dong
author_facet Liu, Rui
Yu, Dian
Ke, Lei
Liu, Haolin
Zhou, Yujun
Liang, Zhenwen
Mi, Haitao
Tokekar, Pratap
Yu, Dong
contents Reinforcement Learning with Verifiable Rewards (RLVR) has become a key paradigm to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevalent group-based algorithms such as GRPO require multi-rollout sampling for each prompt. While more efficient single-rollout variants have recently been explored in text-only settings, we find that they suffer from severe instability in multimodal contexts, often leading to training collapse. To address this training efficiency-stability trade-off, we introduce $\textbf{MSSR}$ (Multimodal Stabilized Single-Rollout), a group-free RLVR framework that achieves both stable optimization and effective multimodal reasoning performance. MSSR achieves this via an entropy-based advantage-shaping mechanism that adaptively regularizes advantage magnitudes, preventing collapse and maintaining training stability. While such mechanisms have been used in group-based RLVR, we show that in the multimodal single-rollout setting they are not merely beneficial but essential for stability. In in-distribution evaluations, MSSR demonstrates superior training compute efficiency, achieving similar validation accuracy to the group-based baseline with half the training steps. When trained for the same number of steps, MSSR's performance surpasses the group-based baseline and shows consistent generalization improvements across five diverse reasoning-intensive benchmarks. Together, these results demonstrate that MSSR enables stable, compute-efficient, and effective RLVR for complex multimodal reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stable and Efficient Single-Rollout RL for Multimodal Reasoning
Liu, Rui
Yu, Dian
Ke, Lei
Liu, Haolin
Zhou, Yujun
Liang, Zhenwen
Mi, Haitao
Tokekar, Pratap
Yu, Dong
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Reinforcement Learning with Verifiable Rewards (RLVR) has become a key paradigm to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevalent group-based algorithms such as GRPO require multi-rollout sampling for each prompt. While more efficient single-rollout variants have recently been explored in text-only settings, we find that they suffer from severe instability in multimodal contexts, often leading to training collapse. To address this training efficiency-stability trade-off, we introduce $\textbf{MSSR}$ (Multimodal Stabilized Single-Rollout), a group-free RLVR framework that achieves both stable optimization and effective multimodal reasoning performance. MSSR achieves this via an entropy-based advantage-shaping mechanism that adaptively regularizes advantage magnitudes, preventing collapse and maintaining training stability. While such mechanisms have been used in group-based RLVR, we show that in the multimodal single-rollout setting they are not merely beneficial but essential for stability. In in-distribution evaluations, MSSR demonstrates superior training compute efficiency, achieving similar validation accuracy to the group-based baseline with half the training steps. When trained for the same number of steps, MSSR's performance surpasses the group-based baseline and shows consistent generalization improvements across five diverse reasoning-intensive benchmarks. Together, these results demonstrate that MSSR enables stable, compute-efficient, and effective RLVR for complex multimodal reasoning tasks.
title Stable and Efficient Single-Rollout RL for Multimodal Reasoning
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.18215