CoPRIS: Efficient and Stable Reinforcement Learning via Concurrency-Controlled Partial Rollout with Importance Sampling
Fuente:
arXiv
Guardado en:
| Autores principales: | Qu, Zekai, Pan, Yinxu, Sun, Ao, Xiao, Chaojun, Han, Xu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
por: Xu, Yixuan Even, et al.
Publicado: (2025)
por: Xu, Yixuan Even, et al.
Publicado: (2025)
APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
por: Zhou, Yuzhen, et al.
Publicado: (2025)
por: Zhou, Yuzhen, et al.
Publicado: (2025)
Partial Domain Adaptation via Importance Sampling-based Shift Correction
por: Guo, Cheng-Jun, et al.
Publicado: (2025)
por: Guo, Cheng-Jun, et al.
Publicado: (2025)
EchoRL: Reinforcement Learning via Rollout Echoing
por: Bi, Jinhe, et al.
Publicado: (2026)
por: Bi, Jinhe, et al.
Publicado: (2026)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
por: Luo, Sijia, et al.
Publicado: (2026)
por: Luo, Sijia, et al.
Publicado: (2026)
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
por: Wang, Haoran, et al.
Publicado: (2023)
por: Wang, Haoran, et al.
Publicado: (2023)
RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
por: Li, Gang, et al.
Publicado: (2025)
por: Li, Gang, et al.
Publicado: (2025)
Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning
por: Zhang, Zhi, et al.
Publicado: (2026)
por: Zhang, Zhi, et al.
Publicado: (2026)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
por: Zheng, Haizhong, et al.
Publicado: (2025)
por: Zheng, Haizhong, et al.
Publicado: (2025)
Portfolio Reinforcement Learning with Scenario-Context Rollout
por: Bendatu, Vanya Priscillia, et al.
Publicado: (2026)
por: Bendatu, Vanya Priscillia, et al.
Publicado: (2026)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
por: Liu, Bingshuai, et al.
Publicado: (2025)
por: Liu, Bingshuai, et al.
Publicado: (2025)
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
por: Li, Yitong, et al.
Publicado: (2026)
por: Li, Yitong, et al.
Publicado: (2026)
Superior Computer Chess with Model Predictive Control, Reinforcement Learning, and Rollout
por: Gundawar, Atharva, et al.
Publicado: (2024)
por: Gundawar, Atharva, et al.
Publicado: (2024)
More Efficient Randomized Exploration for Reinforcement Learning via Approximate Sampling
por: Ishfaq, Haque, et al.
Publicado: (2024)
por: Ishfaq, Haque, et al.
Publicado: (2024)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
por: Liu, Rui, et al.
Publicado: (2025)
por: Liu, Rui, et al.
Publicado: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
por: Lu, Xiaodong, et al.
Publicado: (2026)
por: Lu, Xiaodong, et al.
Publicado: (2026)
SrSv: Integrating Sequential Rollouts with Sequential Value Estimation for Multi-agent Reinforcement Learning
por: Wan, Xu, et al.
Publicado: (2025)
por: Wan, Xu, et al.
Publicado: (2025)
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
por: Gao, Cheng, et al.
Publicado: (2026)
por: Gao, Cheng, et al.
Publicado: (2026)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
por: Zhang, Hongming, et al.
Publicado: (2023)
por: Zhang, Hongming, et al.
Publicado: (2023)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
por: Heuillet, Maxime, et al.
Publicado: (2025)
por: Heuillet, Maxime, et al.
Publicado: (2025)
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
por: Shen, Jianghan, et al.
Publicado: (2026)
por: Shen, Jianghan, et al.
Publicado: (2026)
Reinforcement Learning for Control with Probabilistic Stability Guarantee: A Finite-Sample Approach
por: Han, Minghao, et al.
Publicado: (2026)
por: Han, Minghao, et al.
Publicado: (2026)
PyBench: Evaluating LLM Agent on various real-world coding tasks
por: Zhang, Yaolun, et al.
Publicado: (2024)
por: Zhang, Yaolun, et al.
Publicado: (2024)
The Elephant in the Room: Rethinking the Usage of Pre-trained Language Model in Sequential Recommendation
por: Qu, Zekai, et al.
Publicado: (2024)
por: Qu, Zekai, et al.
Publicado: (2024)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
por: Nguyen, Hieu Trung, et al.
Publicado: (2026)
por: Nguyen, Hieu Trung, et al.
Publicado: (2026)
BQSched: A Non-intrusive Scheduler for Batch Concurrent Queries via Reinforcement Learning
por: Xu, Chenhao, et al.
Publicado: (2025)
por: Xu, Chenhao, et al.
Publicado: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
por: Xiao, Chaojun, et al.
Publicado: (2024)
por: Xiao, Chaojun, et al.
Publicado: (2024)
Improving Execution Concurrency in Partial-Order Plans via Block-Substitution
por: Noor, Sabah Binte, et al.
Publicado: (2024)
por: Noor, Sabah Binte, et al.
Publicado: (2024)
Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
por: Nie, Chaojun, et al.
Publicado: (2025)
por: Nie, Chaojun, et al.
Publicado: (2025)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
por: Pan, Rui, et al.
Publicado: (2024)
por: Pan, Rui, et al.
Publicado: (2024)
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
por: Xu, Yuhang, et al.
Publicado: (2026)
por: Xu, Yuhang, et al.
Publicado: (2026)
Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning
por: Humayoo, Mahammad, et al.
Publicado: (2018)
por: Humayoo, Mahammad, et al.
Publicado: (2018)
Learning Rollout from Sampling:An R1-Style Tokenized Traffic Simulation Model
por: Wang, Ziyan, et al.
Publicado: (2026)
por: Wang, Ziyan, et al.
Publicado: (2026)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
por: Huang, Yuxiang, et al.
Publicado: (2026)
por: Huang, Yuxiang, et al.
Publicado: (2026)
Multiple Unmanned Aerial Vehicle Formation Control through Deep Reinforcement Learning with Offline Sample Correction
por: Zhongkai Chen, et al.
Publicado: (2025)
por: Zhongkai Chen, et al.
Publicado: (2025)
RTMC: Step-Level Credit Assignment via Rollout Trees
por: Wang, Tao, et al.
Publicado: (2026)
por: Wang, Tao, et al.
Publicado: (2026)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
por: Pang, Jing-Cheng, et al.
Publicado: (2024)
Stable Reinforcement Learning for Efficient Reasoning
por: Dai, Muzhi, et al.
Publicado: (2025)
por: Dai, Muzhi, et al.
Publicado: (2025)
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
por: Wang, Yujie, et al.
Publicado: (2026)
por: Wang, Yujie, et al.
Publicado: (2026)
SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout
por: Jiang, Chiyu Max, et al.
Publicado: (2024)
por: Jiang, Chiyu Max, et al.
Publicado: (2024)
Ejemplares similares
-
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
por: Xu, Yixuan Even, et al.
Publicado: (2025) -
APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
por: Zhou, Yuzhen, et al.
Publicado: (2025) -
Partial Domain Adaptation via Importance Sampling-based Shift Correction
por: Guo, Cheng-Jun, et al.
Publicado: (2025) -
EchoRL: Reinforcement Learning via Rollout Echoing
por: Bi, Jinhe, et al.
Publicado: (2026) -
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
por: Luo, Sijia, et al.
Publicado: (2026)