Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Tianshu, Zhang, Wenyu, Zuo, Xiaoying, Tian, Lun, Zhao, Haotian, Zeng, Yucheng, Gu, Jingnan, Dong, Daxiang, Wu, Jianmin, Yin, Dawei, Shen, Dou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning
von: Zhao, Haotian, et al.
Veröffentlicht: (2026)
von: Zhao, Haotian, et al.
Veröffentlicht: (2026)
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
von: Zhang, Wenyu, et al.
Veröffentlicht: (2026)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2026)
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
von: Dong, Daxiang, et al.
Veröffentlicht: (2026)
von: Dong, Daxiang, et al.
Veröffentlicht: (2026)
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
von: Zeng, Yucheng, et al.
Veröffentlicht: (2026)
von: Zeng, Yucheng, et al.
Veröffentlicht: (2026)
QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs
von: Li, Shupeng, et al.
Veröffentlicht: (2025)
von: Li, Shupeng, et al.
Veröffentlicht: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
ROAST: Rollout-based On-distribution Activation Steering Technique
von: Su, Xuanbo, et al.
Veröffentlicht: (2026)
von: Su, Xuanbo, et al.
Veröffentlicht: (2026)
Agentic Flow Steering and Parallel Rollout Search for Spatially Grounded Text-to-Image Generation
von: Chen, Ping, et al.
Veröffentlicht: (2026)
von: Chen, Ping, et al.
Veröffentlicht: (2026)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
EchoRL: Reinforcement Learning via Rollout Echoing
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
von: Gu, Hao, et al.
Veröffentlicht: (2026)
von: Gu, Hao, et al.
Veröffentlicht: (2026)
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
Heddle: A Distributed Orchestration System for Agentic RL Rollout
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
von: Djuhera, Aladin, et al.
Veröffentlicht: (2026)
von: Djuhera, Aladin, et al.
Veröffentlicht: (2026)
SteerRM: Debiasing Reward Models via Sparse Autoencoders
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
von: Sun, Mengyuan, et al.
Veröffentlicht: (2026)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
von: Zhang, Xuechen, et al.
Veröffentlicht: (2025)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2025)
Towards Better RL Training Data Utilization via Second-Order Rollout
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
von: Gu, Chengyang, et al.
Veröffentlicht: (2026)
von: Gu, Chengyang, et al.
Veröffentlicht: (2026)
Most Likely Sequence Generation for $n$-Grams, Transformers, HMMs, and Markov Chains, by Using Rollout Algorithms
von: Li, Yuchao, et al.
Veröffentlicht: (2024)
von: Li, Yuchao, et al.
Veröffentlicht: (2024)
Autoequivalences of Derived Categories of Moduli Spaces of Vector Bundles
von: Zuo, Haotian
Veröffentlicht: (2026)
von: Zuo, Haotian
Veröffentlicht: (2026)
Actions Speak Louder Than Words: Rate-Reward Trade-off in Markov Decision Processes
von: Wu, Haotian, et al.
Veröffentlicht: (2025)
von: Wu, Haotian, et al.
Veröffentlicht: (2025)
Joint Training Across Multiple Activation Sparsity Regimes
von: Wang, Haotian
Veröffentlicht: (2026)
von: Wang, Haotian
Veröffentlicht: (2026)
A Rollout-Based Algorithm and Reward Function for Resource Allocation in Business Processes
von: Middelhuis, Jeroen, et al.
Veröffentlicht: (2025)
von: Middelhuis, Jeroen, et al.
Veröffentlicht: (2025)
Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
von: Gao, Wei, et al.
Veröffentlicht: (2025)
von: Gao, Wei, et al.
Veröffentlicht: (2025)
When Less Is More: Binary Feedback Can Outperform Ordinal Comparisons in Ranking Recovery
von: Xu, Shirong, et al.
Veröffentlicht: (2025)
von: Xu, Shirong, et al.
Veröffentlicht: (2025)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Long-Run Average Reward Maximization of A Regulated Regime-Switching Diffusion Model
von: Zeng, Lingjia, et al.
Veröffentlicht: (2025)
von: Zeng, Lingjia, et al.
Veröffentlicht: (2025)
Graph Fractional Hilbert Transform: Theory and Application
von: Li, Daxiang, et al.
Veröffentlicht: (2025)
von: Li, Daxiang, et al.
Veröffentlicht: (2025)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
von: Muslimani, Calarina, et al.
Veröffentlicht: (2025)
von: Muslimani, Calarina, et al.
Veröffentlicht: (2025)
Molecular Basis of Aggressiveness in Pituitary Adenomas and Its Association With the Immune Microenvironment
von: Xiaoyan Chen, et al.
Veröffentlicht: (2025)
von: Xiaoyan Chen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning
von: Zhao, Haotian, et al.
Veröffentlicht: (2026) -
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
von: Zhang, Wenyu, et al.
Veröffentlicht: (2026) -
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
von: Dong, Daxiang, et al.
Veröffentlicht: (2026) -
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
von: Zeng, Yucheng, et al.
Veröffentlicht: (2026) -
QianfanHuijin Technical Report: A Novel Multi-Stage Training Paradigm for Finance Industrial LLMs
von: Li, Shupeng, et al.
Veröffentlicht: (2025)