RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Gang, Qin, Yulei, Tan, Xiaoyu, Yang, Dingkang, Shi, Yuchen, Xu, Zihan, Li, Xiang, Sun, Xing, Li, Ke |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
by: Qin, Yulei, et al.
Published: (2025)
by: Qin, Yulei, et al.
Published: (2025)
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
by: Lin, Haojia, et al.
Published: (2025)
by: Lin, Haojia, et al.
Published: (2025)
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
Training-Free Group Relative Policy Optimization
by: Cai, Yuzheng, et al.
Published: (2025)
by: Cai, Yuzheng, et al.
Published: (2025)
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
by: Qin, Yulei, et al.
Published: (2025)
by: Qin, Yulei, et al.
Published: (2025)
Unleashing the Power of Data Tsunami: A Comprehensive Survey on Data Assessment and Selection for Instruction Tuning of Language Models
by: Qin, Yulei, et al.
Published: (2024)
by: Qin, Yulei, et al.
Published: (2024)
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
by: Lin, Liuhao, et al.
Published: (2025)
by: Lin, Liuhao, et al.
Published: (2025)
Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic Rewards
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models
by: Yang, Yuncheng, et al.
Published: (2024)
by: Yang, Yuncheng, et al.
Published: (2024)
Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization
by: Shi, Yuchen, et al.
Published: (2025)
by: Shi, Yuchen, et al.
Published: (2025)
FlowAgent: Achieving Compliance and Flexibility for Workflow Agents
by: Shi, Yuchen, et al.
Published: (2025)
by: Shi, Yuchen, et al.
Published: (2025)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards
by: Xing, Shangyu, et al.
Published: (2025)
by: Xing, Shangyu, et al.
Published: (2025)
Enhanced RoBERTaSN Model for Industrial IoT Text Similarity Analysis in Smart Manufacturing Systems
by: Maochun Xu, et al.
Published: (2025)
by: Maochun Xu, et al.
Published: (2025)
VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
by: Wang, Yiting, et al.
Published: (2025)
by: Wang, Yiting, et al.
Published: (2025)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
by: Zhou, Qinhao, et al.
Published: (2024)
by: Zhou, Qinhao, et al.
Published: (2024)
Sinkhorn Distance Minimization for Knowledge Distillation
by: Cui, Xiao, et al.
Published: (2024)
by: Cui, Xiao, et al.
Published: (2024)
NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
by: Liu, Xiangyan, et al.
Published: (2025)
by: Liu, Xiangyan, et al.
Published: (2025)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Video Spatial Reasoning with Object-Centric 3D Rollout
by: Tang, Haoran, et al.
Published: (2025)
by: Tang, Haoran, et al.
Published: (2025)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
by: Li, Yuhang, et al.
Published: (2026)
by: Li, Yuhang, et al.
Published: (2026)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
by: Luo, Sijia, et al.
Published: (2026)
by: Luo, Sijia, et al.
Published: (2026)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
by: Wang, Tao, et al.
Published: (2026)
by: Wang, Tao, et al.
Published: (2026)
Light‐Enhanced Tandem‐Responsive Nano Delivery Platform for Amplified Anti‐tumor Efficiency
by: Xing Wang, et al.
Published: (2024)
by: Xing Wang, et al.
Published: (2024)
Superior Computer Chess with Model Predictive Control, Reinforcement Learning, and Rollout
by: Gundawar, Atharva, et al.
Published: (2024)
by: Gundawar, Atharva, et al.
Published: (2024)
Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
by: Shen, Yunzhe, et al.
Published: (2025)
by: Shen, Yunzhe, et al.
Published: (2025)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
by: Ding, Zihan, et al.
Published: (2024)
by: Ding, Zihan, et al.
Published: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
by: Mundada, Gagan, et al.
Published: (2026)
by: Mundada, Gagan, et al.
Published: (2026)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Modular Embedding Recomposition for Incremental Learning
by: Panariello, Aniello, et al.
Published: (2025)
by: Panariello, Aniello, et al.
Published: (2025)
Dr. Assistant: Enhancing Clinical Diagnostic Inquiry via Structured Diagnostic Reasoning Data and Reinforcement Learning
by: Guo, Yue, et al.
Published: (2026)
by: Guo, Yue, et al.
Published: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
EchoRL: Reinforcement Learning via Rollout Echoing
by: Bi, Jinhe, et al.
Published: (2026)
by: Bi, Jinhe, et al.
Published: (2026)
CIDGMed: Causal Inference-Driven Medication Recommendation with Enhanced Dual-Granularity Learning
by: Liang, Shunpan, et al.
Published: (2024)
by: Liang, Shunpan, et al.
Published: (2024)
Exploring Fine-Grained Representation and Recomposition for Cloth-Changing Person Re-Identification
by: Wang, Qizao, et al.
Published: (2023)
by: Wang, Qizao, et al.
Published: (2023)
Similar Items
-
Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models
by: Qin, Yulei, et al.
Published: (2025) -
CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
by: Lin, Haojia, et al.
Published: (2025) -
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
by: Cai, Shaofei, et al.
Published: (2025) -
Training-Free Group Relative Policy Optimization
by: Cai, Yuzheng, et al.
Published: (2025) -
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
by: Qin, Yulei, et al.
Published: (2025)