How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Yangyi, Lin, Jiaye, Fu, Xiaoliang, Qin, Cong, Shi, Haolin, Hu, Chaowen, Pan, Lu, Zeng, Ke, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time Scaling
by: Wang, Xinglin, et al.
Published: (2025)
by: Wang, Xinglin, et al.
Published: (2025)
Policy Optimization for Dynamic Heart Transplant Allocation
by: Anagnostides, Ioannis, et al.
Published: (2025)
by: Anagnostides, Ioannis, et al.
Published: (2025)
Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR
by: Wang, Tao, et al.
Published: (2026)
by: Wang, Tao, et al.
Published: (2026)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
by: Nguyen, Hieu Trung, et al.
Published: (2026)
by: Nguyen, Hieu Trung, et al.
Published: (2026)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
by: Zheng, Binbin, et al.
Published: (2026)
by: Zheng, Binbin, et al.
Published: (2026)
How to Allocate Personnel Costs of Reference.
by: Spencer, Carol
Published: (1974)
by: Spencer, Carol
Published: (1974)
Second Price Matching with Complete Allocation and Degree Constraints
by: Pinchasi, Rom, et al.
Published: (2025)
by: Pinchasi, Rom, et al.
Published: (2025)
A Rollout-Based Algorithm and Reward Function for Resource Allocation in Business Processes
by: Middelhuis, Jeroen, et al.
Published: (2025)
by: Middelhuis, Jeroen, et al.
Published: (2025)
Autoregressive Policy Optimization for Constrained Allocation Tasks
by: Winkel, David, et al.
Published: (2024)
by: Winkel, David, et al.
Published: (2024)
Allocation, Distribution, and Policy
by: Bowles, Samuel
Published: (2025)
by: Bowles, Samuel
Published: (2025)
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation
by: Shen, Yiqun, et al.
Published: (2025)
by: Shen, Yiqun, et al.
Published: (2025)
How Do Climate‐Related Risks and Opportunities Affect Portfolio Allocation and Asset Pricing?
by: Maher Asal, et al.
Published: (2025)
by: Maher Asal, et al.
Published: (2025)
Consensus-Based Dynamic Task Allocation for Multi-Robot System Considering Payloads Consumption
by: Qiu, Xuekai, et al.
Published: (2024)
by: Qiu, Xuekai, et al.
Published: (2024)
On-Policy Distillation with Best-of-N Teacher Rollout Selection
by: Zhang, Ke, et al.
Published: (2026)
by: Zhang, Ke, et al.
Published: (2026)
Risk-Aware Allocation of Transmission Capacity for AI Data Centers
by: Li, Shaoze, et al.
Published: (2026)
by: Li, Shaoze, et al.
Published: (2026)
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients
by: Xu, Mingwei, et al.
Published: (2026)
by: Xu, Mingwei, et al.
Published: (2026)
Skill or Skip? Learning Selective Skill Invocation in Agentic Tasks via Dual-Granularity Preference Learning
by: Chen, Chishui, et al.
Published: (2026)
by: Chen, Chishui, et al.
Published: (2026)
MAPO: Mixed Advantage Policy Optimization
by: Huang, Wenke, et al.
Published: (2025)
by: Huang, Wenke, et al.
Published: (2025)
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
by: Feldman, Shai, et al.
Published: (2026)
by: Feldman, Shai, et al.
Published: (2026)
NHANES‐Derived Machine Learning Model for Early Identification of Frailty Risk in CKM: Optimizing Resource Allocation Through Predictive Analytics
by: Wenlong Ding, et al.
Published: (2026)
by: Wenlong Ding, et al.
Published: (2026)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
by: Yao, Zhiyuan, et al.
Published: (2026)
by: Yao, Zhiyuan, et al.
Published: (2026)
FedSAC: Dynamic Submodel Allocation for Collaborative Fairness in Federated Learning
by: Wang, Zihui, et al.
Published: (2024)
by: Wang, Zihui, et al.
Published: (2024)
AM-PPO: (Advantage) Alpha-Modulation with Proximal Policy Optimization
by: Sane, Soham
Published: (2025)
by: Sane, Soham
Published: (2025)
How Is Policy Attention Allocated in Crisis Recovery? A Qualitative Comparative Analysis of 31 Provincial Governments in China
by: Yu Zhao, et al.
Published: (2025)
by: Yu Zhao, et al.
Published: (2025)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
Are Full Rollouts Necessary for On-Policy Distillation?
by: Zhang, Yaocheng, et al.
Published: (2026)
by: Zhang, Yaocheng, et al.
Published: (2026)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
Public Goods and Public Allocation Policy
Published: (2020)
Published: (2020)
BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning
by: Gong, Shijin, et al.
Published: (2026)
by: Gong, Shijin, et al.
Published: (2026)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Dynamic Bandwidth Allocation for Hybrid Event-RGB Transmission
by: Yang, Pujing, et al.
Published: (2025)
by: Yang, Pujing, et al.
Published: (2025)
Learning Policies for Dynamic Coalition Formation in Multi-Robot Task Allocation
by: Bezerra, Lucas C. D., et al.
Published: (2024)
by: Bezerra, Lucas C. D., et al.
Published: (2024)
AI-Empowered Resource Allocation for Wirelessly Powered Pinching-Antenna Systems
by: Pakravan, Saeid, et al.
Published: (2026)
by: Pakravan, Saeid, et al.
Published: (2026)
Bandit Allocational Instability
by: Chen, Yilun, et al.
Published: (2026)
by: Chen, Yilun, et al.
Published: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
by: Henry, James
Published: (2026)
by: Henry, James
Published: (2026)
Similar Items
-
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
by: Fu, Xiaoliang, et al.
Published: (2026) -
Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning
by: Fang, Yangyi, et al.
Published: (2026) -
From $\log π$ to $π$: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight
by: Fu, Xiaoliang, et al.
Published: (2026) -
Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training
by: Fang, Yangyi, et al.
Published: (2026) -
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)