Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yuheng, Yao, Wenlin, Yu, Changlong, Liu, Yao, Yin, Qingyu, Yin, Bing, Yun, Hyokun, Li, Lihong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025)
by: Wei, Zhepei, et al.
Published: (2025)
Ask a Strong LLM Judge when Your Reward Model is Uncertain
by: Xu, Zhenghao, et al.
Published: (2025)
by: Xu, Zhenghao, et al.
Published: (2025)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
by: Wang, Tao, et al.
Published: (2026)
by: Wang, Tao, et al.
Published: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
by: Miao, Yuchun, et al.
Published: (2026)
by: Miao, Yuchun, et al.
Published: (2026)
Evaluating Parameter Efficient Methods for RLVR
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
HeaPA: Difficulty-Aware Heap Sampling and On-Policy Query Augmentation for LLM Reinforcement Learning
by: Wang, Weiqi, et al.
Published: (2026)
by: Wang, Weiqi, et al.
Published: (2026)
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
by: Corrado, Nicholas E., et al.
Published: (2025)
by: Corrado, Nicholas E., et al.
Published: (2025)
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection
by: Wu, Jianghao, et al.
Published: (2026)
by: Wu, Jianghao, et al.
Published: (2026)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
by: Kim, Soeun, et al.
Published: (2026)
by: Kim, Soeun, et al.
Published: (2026)
Gradient-Informed Temporal Sampling Improves Rollout Accuracy in PDE Surrogate Training
by: Wang, Wenshuo, et al.
Published: (2026)
by: Wang, Wenshuo, et al.
Published: (2026)
Evolutionary Contrastive Distillation for Language Model Alignment
by: Katz-Samuels, Julian, et al.
Published: (2024)
by: Katz-Samuels, Julian, et al.
Published: (2024)
Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
by: Lu, Yining, et al.
Published: (2025)
by: Lu, Yining, et al.
Published: (2025)
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
by: Zhang, Yuheng, et al.
Published: (2026)
by: Zhang, Yuheng, et al.
Published: (2026)
RLVR-World: Training World Models with Reinforcement Learning
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and Routing
by: Li, Conglong, et al.
Published: (2022)
by: Li, Conglong, et al.
Published: (2022)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
RLPR: Extrapolating RLVR to General Domains without Verifiers
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
MBDP: A Model-based Approach to Achieve both Robustness and Sample Efficiency via Double Dropout Planning
by: Zhang, Wanpeng, et al.
Published: (2021)
by: Zhang, Wanpeng, et al.
Published: (2021)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
by: Mundada, Gagan, et al.
Published: (2026)
by: Mundada, Gagan, et al.
Published: (2026)
Improving Sample Efficiency of Model-Free Algorithms for Zero-Sum Markov Games
by: Feng, Songtao, et al.
Published: (2023)
by: Feng, Songtao, et al.
Published: (2023)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Controlling Transient Amplification Improves Long-horizon Rollouts
by: Pervez, Adeel, et al.
Published: (2026)
by: Pervez, Adeel, et al.
Published: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
by: Lu, Xiaodong, et al.
Published: (2026)
by: Lu, Xiaodong, et al.
Published: (2026)
Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent
by: Shu, Yao, et al.
Published: (2026)
by: Shu, Yao, et al.
Published: (2026)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
by: Liu, Bingshuai, et al.
Published: (2025)
by: Liu, Bingshuai, et al.
Published: (2025)
Distribution-Free Fair Federated Learning with Small Samples
by: Yin, Qichuan, et al.
Published: (2024)
by: Yin, Qichuan, et al.
Published: (2024)
LEASE: Offline Preference-based Reinforcement Learning with High Sample Efficiency
by: Liu, Xiao-Yin, et al.
Published: (2024)
by: Liu, Xiao-Yin, et al.
Published: (2024)
The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
Not only where, But when: Temporal Scheduling for RLVR
by: Zhang, Jinghao, et al.
Published: (2026)
by: Zhang, Jinghao, et al.
Published: (2026)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
by: Zhai, Zhiyuan, et al.
Published: (2026)
by: Zhai, Zhiyuan, et al.
Published: (2026)
Towards Trustworthy GUI Agents: A Survey
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
Novelty-based Sample Reuse for Continuous Robotics Control
by: Duan, Ke, et al.
Published: (2024)
by: Duan, Ke, et al.
Published: (2024)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
by: Wang, Xinglin, et al.
Published: (2025)
by: Wang, Xinglin, et al.
Published: (2025)
Improving Sample Efficiency of Reinforcement Learning with Background Knowledge from Large Language Models
by: Zhang, Fuxiang, et al.
Published: (2024)
by: Zhang, Fuxiang, et al.
Published: (2024)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
by: Yue, Murong, et al.
Published: (2024)
by: Yue, Murong, et al.
Published: (2024)
Similar Items
-
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025) -
Ask a Strong LLM Judge when Your Reward Model is Uncertain
by: Xu, Zhenghao, et al.
Published: (2025) -
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
by: Wang, Tao, et al.
Published: (2026) -
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
by: Miao, Yuchun, et al.
Published: (2026) -
Evaluating Parameter Efficient Methods for RLVR
by: Yin, Qingyu, et al.
Published: (2025)