CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Zhiyuan, Zhang, Yi-Kai, Chen, Yuxin, Sun, Yueqing, Xu, Zishan, Yang, Yu, Hu, Tianhao, Gu, Qi, Su, Hui, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
by: Zhang, Yi-Kai, et al.
Published: (2026)
by: Zhang, Yi-Kai, et al.
Published: (2026)
$V_0$: A Generalist Value Model for Any Policy at State Zero
by: Zhang, Yi-Kai, et al.
Published: (2026)
by: Zhang, Yi-Kai, et al.
Published: (2026)
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
by: Jin, Kyohoon, et al.
Published: (2025)
by: Jin, Kyohoon, et al.
Published: (2025)
CoBA: Integrated Deep Learning Model for Reliable Low-Altitude UAV Classification in mmWave Radio Networks
by: Sajid, Junaid, et al.
Published: (2026)
by: Sajid, Junaid, et al.
Published: (2026)
TopoCurate:Modeling Interaction Topology for Tool-Use Agent Training
by: Yang, Jinluan, et al.
Published: (2026)
by: Yang, Jinluan, et al.
Published: (2026)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
by: Li, Ziniu, et al.
Published: (2025)
by: Li, Ziniu, et al.
Published: (2025)
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
by: Lv, Minxuan, et al.
Published: (2026)
by: Lv, Minxuan, et al.
Published: (2026)
Effort as Ceiling, Not Dial: Reasoning Budget Does Not Modulate Cognitive Cost Alignment Between Humans and Large Reasoning Models
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning
by: Li, Lin, et al.
Published: (2025)
by: Li, Lin, et al.
Published: (2025)
Budgeting Counterfactual for Offline RL
by: Liu, Yao, et al.
Published: (2023)
by: Liu, Yao, et al.
Published: (2023)
ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training
by: Tu, Dunwei, et al.
Published: (2026)
by: Tu, Dunwei, et al.
Published: (2026)
SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
by: Lu, Zhengxi, et al.
Published: (2026)
by: Lu, Zhengxi, et al.
Published: (2026)
Engineering-Oriented Symbolic Regression: LLMs as Physics Agents for Discovery of Simulation-Ready Constitutive Laws
by: Wu, Yue, et al.
Published: (2026)
by: Wu, Yue, et al.
Published: (2026)
MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoning
by: Liu, Yuxin, et al.
Published: (2026)
by: Liu, Yuxin, et al.
Published: (2026)
DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs
by: Song, Mingxuan, et al.
Published: (2026)
by: Song, Mingxuan, et al.
Published: (2026)
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
by: Guo, Yifu, et al.
Published: (2025)
by: Guo, Yifu, et al.
Published: (2025)
STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
by: Gu, Chengyang, et al.
Published: (2026)
by: Gu, Chengyang, et al.
Published: (2026)
DAST: Context-Aware Compression in LLMs via Dynamic Allocation of Soft Tokens
by: Chen, Shaoshen, et al.
Published: (2025)
by: Chen, Shaoshen, et al.
Published: (2025)
DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
by: Hu, Haoyu, et al.
Published: (2026)
by: Hu, Haoyu, et al.
Published: (2026)
Self-Distilled Agentic Reinforcement Learning
by: Lu, Zhengxi, et al.
Published: (2026)
by: Lu, Zhengxi, et al.
Published: (2026)
ObjectRL: An Object-Oriented Reinforcement Learning Codebase
by: Baykal, Gulcin, et al.
Published: (2025)
by: Baykal, Gulcin, et al.
Published: (2025)
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
by: Hu, Tianhao, et al.
Published: (2026)
by: Hu, Tianhao, et al.
Published: (2026)
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
by: Fan, Zhiyuan, et al.
Published: (2026)
by: Fan, Zhiyuan, et al.
Published: (2026)
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
Unraveling the Mystery of Scaling Laws: Part I
by: Su, Hui, et al.
Published: (2024)
by: Su, Hui, et al.
Published: (2024)
Learning to Self-Verify Makes Language Models Better Reasoners
by: Chen, Yuxin, et al.
Published: (2026)
by: Chen, Yuxin, et al.
Published: (2026)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
by: Shi, Wentao, et al.
Published: (2026)
by: Shi, Wentao, et al.
Published: (2026)
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
by: Dong, Yihong, et al.
Published: (2025)
by: Dong, Yihong, et al.
Published: (2025)
Mono2Sls: Automated Monolith-to-Serverless Migration via Multi-Stage Pipeline with Static Analysis
by: Chen, Xingyan, et al.
Published: (2026)
by: Chen, Xingyan, et al.
Published: (2026)
A Budget-Adaptive Allocation Rule for Optimal Computing Budget Allocation
by: Cao, Zirui, et al.
Published: (2023)
by: Cao, Zirui, et al.
Published: (2023)
Capability-Guided Compression: Toward Interpretability-Aware Budget Allocation for Large Language Models
by: Gupta, Rishaank
Published: (2026)
by: Gupta, Rishaank
Published: (2026)
Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning
by: Shi, Yaorui, et al.
Published: (2026)
by: Shi, Yaorui, et al.
Published: (2026)
LLMs for Resource Allocation: A Participatory Budgeting Approach to Inferring Preferences
by: Damle, Sankarshan, et al.
Published: (2025)
by: Damle, Sankarshan, et al.
Published: (2025)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
by: Wang, Ruipeng, et al.
Published: (2026)
by: Wang, Ruipeng, et al.
Published: (2026)
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
by: Zhao, Weikang, et al.
Published: (2025)
by: Zhao, Weikang, et al.
Published: (2025)
Optimal Perturbation Budget Allocation for Data Poisoning in Offline Reinforcement Learning
by: Qiu, Junnan, et al.
Published: (2025)
by: Qiu, Junnan, et al.
Published: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
by: Pignatelli, Eduardo, et al.
Published: (2024)
by: Pignatelli, Eduardo, et al.
Published: (2024)
Similar Items
-
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
by: Zhang, Yi-Kai, et al.
Published: (2026) -
$V_0$: A Generalist Value Model for Any Policy at State Zero
by: Zhang, Yi-Kai, et al.
Published: (2026) -
CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
by: Jin, Kyohoon, et al.
Published: (2025) -
CoBA: Integrated Deep Learning Model for Reliable Low-Altitude UAV Classification in mmWave Radio Networks
by: Sajid, Junaid, et al.
Published: (2026) -
TopoCurate:Modeling Interaction Topology for Tool-Use Agent Training
by: Yang, Jinluan, et al.
Published: (2026)