Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Linghao, Guan, Yiran, Liang, Dingkang, Ju, Jianzhong, Luo, Zhenbo, Qin, Bin, Luan, Jian, Liu, Yuliang, Bai, Xiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918367005245440
author Zhu, Linghao
Guan, Yiran
Liang, Dingkang
Ju, Jianzhong
Luo, Zhenbo
Qin, Bin
Luan, Jian
Liu, Yuliang
Bai, Xiang
author_facet Zhu, Linghao
Guan, Yiran
Liang, Dingkang
Ju, Jianzhong
Luo, Zhenbo
Qin, Bin
Luan, Jian
Liu, Yuliang
Bai, Xiang
contents Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where most advantages in a batch concentrate near zero, and Rollout Silencing, where the proportion of rollouts contributing non-zero gradients diminishes over time. These issues lead to suboptimal gradient updates and hinder long-term learning efficiency. To address these issues, we propose Shuffle-R1, a simple yet principled framework that improves RL fine-tuning efficiency by dynamically restructuring trajectory sampling and batch composition. It introduces (1) Pairwise Trajectory Sampling, which selects high-contrast trajectories with large advantages to improve gradient signal quality, and (2) Advantage-based Trajectory Shuffle, which increases exposure of valuable rollouts through informed batch reshuffling. Experiments across multiple reasoning benchmarks show that our framework consistently outperforms strong RL baselines with minimal overhead. These results highlight the importance of data-centric adaptations for more efficient RL training in MLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05612
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
Zhu, Linghao
Guan, Yiran
Liang, Dingkang
Ju, Jianzhong
Luo, Zhenbo
Qin, Bin
Luan, Jian
Liu, Yuliang
Bai, Xiang
Machine Learning
Artificial Intelligence
Reinforcement learning (RL) has emerged as an effective post-training paradigm for enhancing the reasoning capabilities of multimodal large language model (MLLM). However, current RL pipelines often suffer from training inefficiencies caused by two underexplored issues: Advantage Collapsing, where most advantages in a batch concentrate near zero, and Rollout Silencing, where the proportion of rollouts contributing non-zero gradients diminishes over time. These issues lead to suboptimal gradient updates and hinder long-term learning efficiency. To address these issues, we propose Shuffle-R1, a simple yet principled framework that improves RL fine-tuning efficiency by dynamically restructuring trajectory sampling and batch composition. It introduces (1) Pairwise Trajectory Sampling, which selects high-contrast trajectories with large advantages to improve gradient signal quality, and (2) Advantage-based Trajectory Shuffle, which increases exposure of valuable rollouts through informed batch reshuffling. Experiments across multiple reasoning benchmarks show that our framework consistently outperforms strong RL baselines with minimal overhead. These results highlight the importance of data-centric adaptations for more efficient RL training in MLLM.
title Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.05612