CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Zhiyuan, Zhang, Yi-Kai, Chen, Yuxin, Sun, Yueqing, Xu, Zishan, Yang, Yu, Hu, Tianhao, Gu, Qi, Su, Hui, Cai, Xunliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914310325796864
author Yao, Zhiyuan
Zhang, Yi-Kai
Chen, Yuxin
Sun, Yueqing
Xu, Zishan
Yang, Yu
Hu, Tianhao
Gu, Qi
Su, Hui
Cai, Xunliang
author_facet Yao, Zhiyuan
Zhang, Yi-Kai
Chen, Yuxin
Sun, Yueqing
Xu, Zishan
Yang, Yu
Hu, Tianhao
Gu, Qi
Su, Hui
Cai, Xunliang
contents Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key approach for enhancing LLM reasoning. However, standard frameworks like Group Relative Policy Optimization (GRPO) typically employ a uniform rollout budget, leading to resource inefficiency. Moreover, existing adaptive methods often rely on instance-level metrics, such as task pass rates, failing to capture the model's dynamic learning state. To address these limitations, we propose CoBA-RL, a reinforcement learning algorithm designed to adaptively allocate rollout budgets based on the model's evolving capability. Specifically, CoBA-RL utilizes a Capability-Oriented Value function to map tasks to their potential training gains and employs a heap-based greedy strategy to efficiently self-calibrate the distribution of computational resources to samples with high training value. Extensive experiments demonstrate that our approach effectively orchestrates the trade-off between exploration and exploitation, delivering consistent generalization improvements across multiple challenging benchmarks. These findings underscore that quantifying sample training value and optimizing budget allocation are pivotal for advancing LLM post-training efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03048
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
Yao, Zhiyuan
Zhang, Yi-Kai
Chen, Yuxin
Sun, Yueqing
Xu, Zishan
Yang, Yu
Hu, Tianhao
Gu, Qi
Su, Hui
Cai, Xunliang
Machine Learning
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key approach for enhancing LLM reasoning. However, standard frameworks like Group Relative Policy Optimization (GRPO) typically employ a uniform rollout budget, leading to resource inefficiency. Moreover, existing adaptive methods often rely on instance-level metrics, such as task pass rates, failing to capture the model's dynamic learning state. To address these limitations, we propose CoBA-RL, a reinforcement learning algorithm designed to adaptively allocate rollout budgets based on the model's evolving capability. Specifically, CoBA-RL utilizes a Capability-Oriented Value function to map tasks to their potential training gains and employs a heap-based greedy strategy to efficiently self-calibrate the distribution of computational resources to samples with high training value. Extensive experiments demonstrate that our approach effectively orchestrates the trade-off between exploration and exploitation, delivering consistent generalization improvements across multiple challenging benchmarks. These findings underscore that quantifying sample training value and optimizing budget allocation are pivotal for advancing LLM post-training efficiency.
title CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03048