From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Donglai, Yang, Hongzheng, Zhao, Yuzhi, Zhang, Pingping, Chen, Jinpeng, Ma, Wenao, Hou, Zhijian, Wu, Mengyang, Li, Xiaolei, Hu, Senkang, Guan, Ziyi, Li, Jason Chun Lok, Po, Lai Man |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
by: Guan, Ziyi, et al.
Published: (2025)
by: Guan, Ziyi, et al.
Published: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
by: Xu, Mingjie, et al.
Published: (2025)
by: Xu, Mingjie, et al.
Published: (2025)
AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)
by: Huang, Fanding, et al.
Published: (2025)
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations
by: Xu, Mingjie, et al.
Published: (2024)
by: Xu, Mingjie, et al.
Published: (2024)
The Exploration-Exploitation Dilemma Revisited: An Entropy Perspective
by: Yan, Renye, et al.
Published: (2024)
by: Yan, Renye, et al.
Published: (2024)
On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment
by: Liu, Zhanyu, et al.
Published: (2026)
by: Liu, Zhanyu, et al.
Published: (2026)
Modeling Dual-Exposure Quad-Bayer Patterns for Joint Denoising and Deblurring
by: Zhao, Yuzhi, et al.
Published: (2024)
by: Zhao, Yuzhi, et al.
Published: (2024)
Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective
by: Hao, Zhezheng, et al.
Published: (2025)
by: Hao, Zhezheng, et al.
Published: (2025)
Fault Tolerant Reconfigurable ML Multiprocessor
by: Li, Tangrui, et al.
Published: (2025)
by: Li, Tangrui, et al.
Published: (2025)
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
by: Bamba, Udbhav, et al.
Published: (2025)
by: Bamba, Udbhav, et al.
Published: (2025)
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
by: Huang, Zhuoxu, et al.
Published: (2026)
by: Huang, Zhuoxu, et al.
Published: (2026)
Energy Exploration & Exploitation
Published: (2020)
Published: (2020)
SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
by: Chen, Jinpeng, et al.
Published: (2025)
by: Chen, Jinpeng, et al.
Published: (2025)
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization
by: Xu, Huimin, et al.
Published: (2026)
by: Xu, Huimin, et al.
Published: (2026)
Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation
by: Lei, Xiaohan, et al.
Published: (2024)
by: Lei, Xiaohan, et al.
Published: (2024)
Disentangling Exploration from Exploitation
by: Lizzeri, Alessandro, et al.
Published: (2024)
by: Lizzeri, Alessandro, et al.
Published: (2024)
Marine Exploration and Exploitation of Hydrocarbons
by: Radovich, Violeta S.
Published: (2025)
by: Radovich, Violeta S.
Published: (2025)
Organizational Factors for Exploration and Exploitation
by: Sharadindu Pandey
Published: (2009)
by: Sharadindu Pandey
Published: (2009)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
by: Gu, Hengrui, et al.
Published: (2026)
by: Gu, Hengrui, et al.
Published: (2026)
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
by: Liu, Jia, et al.
Published: (2025)
by: Liu, Jia, et al.
Published: (2025)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
by: Nguyen, Hieu Trung, et al.
Published: (2026)
by: Nguyen, Hieu Trung, et al.
Published: (2026)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
Skill-Conditioned Gated Self-Distillation for LLM Reasoning
by: Huang, Jiazhen, et al.
Published: (2026)
by: Huang, Jiazhen, et al.
Published: (2026)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
by: Chen, Kun, et al.
Published: (2026)
by: Chen, Kun, et al.
Published: (2026)
Neural Exploitation and Exploration of Contextual Bandits
by: Ban, Yikun, et al.
Published: (2023)
by: Ban, Yikun, et al.
Published: (2023)
In-context Exploration-Exploitation for Reinforcement Learning
by: Dai, Zhenwen, et al.
Published: (2024)
by: Dai, Zhenwen, et al.
Published: (2024)
Exploitation Is All You Need... for Exploration
by: Rentschler, Micah, et al.
Published: (2025)
by: Rentschler, Micah, et al.
Published: (2025)
Exploration, Exploitation, and Organizational Coordination Mechanisms
by: Silvio Popadiuk
Published: (2016)
by: Silvio Popadiuk
Published: (2016)
Adaptive Noise-Tolerant Network for Image Segmentation
by: Li, Weizhi
Published: (2025)
by: Li, Weizhi
Published: (2025)
3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
by: Gong, Shizhan, et al.
Published: (2023)
by: Gong, Shizhan, et al.
Published: (2023)
SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
by: Cui, Wendi, et al.
Published: (2024)
by: Cui, Wendi, et al.
Published: (2024)
Dual Control of Exploration and Exploitation for Auto-Optimisation Control with Active Learning
by: Li, Zhongguo, et al.
Published: (2023)
by: Li, Zhongguo, et al.
Published: (2023)
MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization
by: Zhang, Yiyi, et al.
Published: (2025)
by: Zhang, Yiyi, et al.
Published: (2025)
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
by: Lee, Chanuk, et al.
Published: (2026)
by: Lee, Chanuk, et al.
Published: (2026)
Similar Items
-
KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
by: Guan, Ziyi, et al.
Published: (2025) -
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025) -
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
by: Xu, Mingjie, et al.
Published: (2025) -
AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
by: Li, Kun, et al.
Published: (2025) -
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning
by: Huang, Fanding, et al.
Published: (2025)