From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Donglai, Yang, Hongzheng, Zhao, Yuzhi, Zhang, Pingping, Chen, Jinpeng, Ma, Wenao, Hou, Zhijian, Wu, Mengyang, Li, Xiaolei, Hu, Senkang, Guan, Ziyi, Li, Jason Chun Lok, Po, Lai Man
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917365274378240
author Xu, Donglai
Yang, Hongzheng
Zhao, Yuzhi
Zhang, Pingping
Chen, Jinpeng
Ma, Wenao
Hou, Zhijian
Wu, Mengyang
Li, Xiaolei
Hu, Senkang
Guan, Ziyi
Li, Jason Chun Lok
Po, Lai Man
author_facet Xu, Donglai
Yang, Hongzheng
Zhao, Yuzhi
Zhang, Pingping
Chen, Jinpeng
Ma, Wenao
Hou, Zhijian
Wu, Mengyang
Li, Xiaolei
Hu, Senkang
Guan, Ziyi
Li, Jason Chun Lok
Po, Lai Man
contents Reinforcement Learning with Verifiable Rewards (RLVR) for Multimodal Large Language Models (MLLMs) is highly dependent on high-quality labeled data, which is often scarce and prone to substantial annotation noise in real-world scenarios. Existing unsupervised RLVR methods, including pure entropy minimization, can overfit to incorrect labels and limit the crucial reward ranking signal for Group-Relative Policy Optimization (GRPO). To address these challenges and enhance noise tolerance, we propose a novel two-stage, token-level entropy optimization method for RLVR. This approach dynamically guides the model from exploration to exploitation during training. In the initial exploration phase, token-level entropy maximization promotes diverse and stochastic output generation, serving as a strong regularizer that prevents premature convergence to noisy labels and ensures sufficient intra-group variation, which enables more reliable reward gradient estimation in GRPO. As training progresses, the method transitions into the exploitation phase, where token-level entropy minimization encourages the model to produce confident and deterministic outputs, thereby consolidating acquired knowledge and refining prediction accuracy. Empirically, across three MLLM backbones - Qwen2-VL-2B, Qwen2-VL-7B, and Qwen2.5-VL-3B - spanning diverse noise settings and multiple tasks, our phased strategy consistently outperforms prior approaches by unifying and enhancing external, internal, and entropy-based methods, delivering robust and superior performance across the board.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07738
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
Xu, Donglai
Yang, Hongzheng
Zhao, Yuzhi
Zhang, Pingping
Chen, Jinpeng
Ma, Wenao
Hou, Zhijian
Wu, Mengyang
Li, Xiaolei
Hu, Senkang
Guan, Ziyi
Li, Jason Chun Lok
Po, Lai Man
Machine Learning
Computer Vision and Pattern Recognition
Reinforcement Learning with Verifiable Rewards (RLVR) for Multimodal Large Language Models (MLLMs) is highly dependent on high-quality labeled data, which is often scarce and prone to substantial annotation noise in real-world scenarios. Existing unsupervised RLVR methods, including pure entropy minimization, can overfit to incorrect labels and limit the crucial reward ranking signal for Group-Relative Policy Optimization (GRPO). To address these challenges and enhance noise tolerance, we propose a novel two-stage, token-level entropy optimization method for RLVR. This approach dynamically guides the model from exploration to exploitation during training. In the initial exploration phase, token-level entropy maximization promotes diverse and stochastic output generation, serving as a strong regularizer that prevents premature convergence to noisy labels and ensures sufficient intra-group variation, which enables more reliable reward gradient estimation in GRPO. As training progresses, the method transitions into the exploitation phase, where token-level entropy minimization encourages the model to produce confident and deterministic outputs, thereby consolidating acquired knowledge and refining prediction accuracy. Empirically, across three MLLM backbones - Qwen2-VL-2B, Qwen2-VL-7B, and Qwen2.5-VL-3B - spanning diverse noise settings and multiple tasks, our phased strategy consistently outperforms prior approaches by unifying and enhancing external, internal, and entropy-based methods, delivering robust and superior performance across the board.
title From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.07738