Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zhi, Han, Zhen, Mavromatis, Costas, Zhu, Qi, Zhang, Yunyi, Guan, Sheng, Wang, Dingmin, Zhou, Xiong, Wang, Shuai, Adeshina, Soji, Ioannidis, Vassilis, Rangwala, Huzefa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
von: Pengmei, Zihan, et al.
Veröffentlicht: (2025)
von: Pengmei, Zihan, et al.
Veröffentlicht: (2025)
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
von: Zhang, Yunyi, et al.
Veröffentlicht: (2026)
von: Zhang, Yunyi, et al.
Veröffentlicht: (2026)
HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases
von: Lee, Meng-Chieh, et al.
Veröffentlicht: (2024)
von: Lee, Meng-Chieh, et al.
Veröffentlicht: (2024)
BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answering
von: Mavromatis, Costas, et al.
Veröffentlicht: (2025)
von: Mavromatis, Costas, et al.
Veröffentlicht: (2025)
Hierarchical Lexical Graph for Enhanced Multi-Hop Retrieval
von: Ghassel, Abdellah, et al.
Veröffentlicht: (2025)
von: Ghassel, Abdellah, et al.
Veröffentlicht: (2025)
SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL
von: Hua, Harper, et al.
Veröffentlicht: (2026)
von: Hua, Harper, et al.
Veröffentlicht: (2026)
BioBridge: Bridging Biomedical Foundation Models via Knowledge Graphs
von: Wang, Zifeng, et al.
Veröffentlicht: (2023)
von: Wang, Zifeng, et al.
Veröffentlicht: (2023)
BoundRL: Efficient Structured Text Segmentation through Reinforced Boundary Generation
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
von: He, Yongjun, et al.
Veröffentlicht: (2025)
von: He, Yongjun, et al.
Veröffentlicht: (2025)
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
von: Liu, Wenpu, et al.
Veröffentlicht: (2026)
von: Liu, Wenpu, et al.
Veröffentlicht: (2026)
NetInfoF Framework: Measuring and Exploiting Network Usable Information
von: Lee, Meng-Chieh, et al.
Veröffentlicht: (2024)
von: Lee, Meng-Chieh, et al.
Veröffentlicht: (2024)
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
von: Ou, Jiefu, et al.
Veröffentlicht: (2026)
von: Ou, Jiefu, et al.
Veröffentlicht: (2026)
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
Pushing the Limits of All-Atom Geometric Graph Neural Networks: Pre-Training, Scaling and Zero-Shot Transfer
von: Pengmei, Zihan, et al.
Veröffentlicht: (2024)
von: Pengmei, Zihan, et al.
Veröffentlicht: (2024)
DeCaf: A Causal Decoupling Framework for OOD Generalization on Node Classification
von: Han, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Han, Xiaoxue, et al.
Veröffentlicht: (2024)
DispaRisk: Auditing Fairness Through Usable Information
von: Vasquez, Jonathan, et al.
Veröffentlicht: (2024)
von: Vasquez, Jonathan, et al.
Veröffentlicht: (2024)
Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
von: Denisov-Blanch, Yegor, et al.
Veröffentlicht: (2026)
von: Denisov-Blanch, Yegor, et al.
Veröffentlicht: (2026)
Protein Structure Tokenization: Benchmarking and New Recipe
von: Yuan, Xinyu, et al.
Veröffentlicht: (2025)
von: Yuan, Xinyu, et al.
Veröffentlicht: (2025)
Rollout-Training Co-Design for Efficient LLM-Based Multi-Agent Reinforcement Learning
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
von: Li, Xuchen, et al.
Veröffentlicht: (2025)
von: Li, Xuchen, et al.
Veröffentlicht: (2025)
Relatron: Automating Relational Machine Learning over Relational Databases
von: Chen, Zhikai, et al.
Veröffentlicht: (2026)
von: Chen, Zhikai, et al.
Veröffentlicht: (2026)
GraphStorm: all-in-one graph machine learning framework for industry applications
von: Zheng, Da, et al.
Veröffentlicht: (2024)
von: Zheng, Da, et al.
Veröffentlicht: (2024)
Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
SemPool: Simple, robust, and interpretable KG pooling for enhancing language models
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
von: Mavromatis, Costas, et al.
Veröffentlicht: (2024)
Offline Learning and Forgetting for Reasoning with Large Language Models
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
von: Ni, Tianwei, et al.
Veröffentlicht: (2025)
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
On Rollouts in Model-Based Reinforcement Learning
von: Frauenknecht, Bernd, et al.
Veröffentlicht: (2025)
von: Frauenknecht, Bernd, et al.
Veröffentlicht: (2025)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
Less is More: A Closer Look at Semantic-based Few-Shot Learning
von: Zhou, Chunpeng, et al.
Veröffentlicht: (2024)
von: Zhou, Chunpeng, et al.
Veröffentlicht: (2024)
EchoRL: Reinforcement Learning via Rollout Echoing
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
Bridging the Training-Inference Gap in LLMs by Leveraging Self-Generated Tokens
von: Cen, Zhepeng, et al.
Veröffentlicht: (2024)
von: Cen, Zhepeng, et al.
Veröffentlicht: (2024)
Portfolio Reinforcement Learning with Scenario-Context Rollout
von: Bendatu, Vanya Priscillia, et al.
Veröffentlicht: (2026)
von: Bendatu, Vanya Priscillia, et al.
Veröffentlicht: (2026)
Say More with Less: Understanding Prompt Learning Behaviors through Gist Compression
von: Li, Xinze, et al.
Veröffentlicht: (2024)
von: Li, Xinze, et al.
Veröffentlicht: (2024)
Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-shot Learning
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2022)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2022)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
von: Gai, Jiading, et al.
Veröffentlicht: (2026)
von: Gai, Jiading, et al.
Veröffentlicht: (2026)
When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
von: He, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
von: Pengmei, Zihan, et al.
Veröffentlicht: (2025) -
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
von: Zhang, Yunyi, et al.
Veröffentlicht: (2026) -
HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases
von: Lee, Meng-Chieh, et al.
Veröffentlicht: (2024) -
BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answering
von: Mavromatis, Costas, et al.
Veröffentlicht: (2025) -
Hierarchical Lexical Graph for Enhanced Multi-Hop Retrieval
von: Ghassel, Abdellah, et al.
Veröffentlicht: (2025)