Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yang, Li, Sunzhu, Liu, Shunyu, Fang, Wenkai, Zhang, Kongcheng, Zhao, Jiale, Yang, Jingwen, Zhou, Yihe, Lv, Jianwei, Zheng, Tongya, Lu, Hengtong, Chen, Wei, Xie, Yan, Song, Mingli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914294096986112
author Zhou, Yang
Li, Sunzhu
Liu, Shunyu
Fang, Wenkai
Zhang, Kongcheng
Zhao, Jiale
Yang, Jingwen
Zhou, Yihe
Lv, Jianwei
Zheng, Tongya
Lu, Hengtong
Chen, Wei
Xie, Yan
Song, Mingli
author_facet Zhou, Yang
Li, Sunzhu
Liu, Shunyu
Fang, Wenkai
Zhang, Kongcheng
Zhao, Jiale
Yang, Jingwen
Zhou, Yihe
Lv, Jianwei
Zheng, Tongya
Lu, Hengtong
Chen, Wei
Xie, Yan
Song, Mingli
contents Recent advances in Large Language Models (LLMs) have underscored the potential of Reinforcement Learning (RL) to facilitate the emergence of reasoning capabilities. Despite the encouraging results, a fundamental dilemma persists as RL improvement relies on learning from high-quality samples, yet the exploration for such samples remains bounded by the inherent limitations of LLMs. This, in effect, creates an undesirable cycle in which what cannot be explored cannot be learned. In this work, we propose Rubric-Scaffolded Reinforcement Learning (RuscaRL), a novel instructional scaffolding framework designed to break the exploration bottleneck for general LLM reasoning. Specifically, RuscaRL introduces checklist-style rubrics as (1) explicit scaffolding for exploration during rollout generation, where different rubrics are provided as external guidance within task instructions to steer diverse high-quality responses. This guidance is gradually decayed over time, encouraging the model to internalize the underlying reasoning patterns; (2) verifiable rewards for exploitation during model training, where we can obtain robust LLM-as-a-Judge scores using rubrics as references, enabling effective RL on general reasoning tasks. Extensive experiments demonstrate the superiority of the proposed RuscaRL across various benchmarks, effectively expanding reasoning boundaries under the Best-of-N evaluation. Our code is available at https://github.com/IANNXANG/RuscaRL.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
Zhou, Yang
Li, Sunzhu
Liu, Shunyu
Fang, Wenkai
Zhang, Kongcheng
Zhao, Jiale
Yang, Jingwen
Zhou, Yihe
Lv, Jianwei
Zheng, Tongya
Lu, Hengtong
Chen, Wei
Xie, Yan
Song, Mingli
Machine Learning
Artificial Intelligence
Recent advances in Large Language Models (LLMs) have underscored the potential of Reinforcement Learning (RL) to facilitate the emergence of reasoning capabilities. Despite the encouraging results, a fundamental dilemma persists as RL improvement relies on learning from high-quality samples, yet the exploration for such samples remains bounded by the inherent limitations of LLMs. This, in effect, creates an undesirable cycle in which what cannot be explored cannot be learned. In this work, we propose Rubric-Scaffolded Reinforcement Learning (RuscaRL), a novel instructional scaffolding framework designed to break the exploration bottleneck for general LLM reasoning. Specifically, RuscaRL introduces checklist-style rubrics as (1) explicit scaffolding for exploration during rollout generation, where different rubrics are provided as external guidance within task instructions to steer diverse high-quality responses. This guidance is gradually decayed over time, encouraging the model to internalize the underlying reasoning patterns; (2) verifiable rewards for exploitation during model training, where we can obtain robust LLM-as-a-Judge scores using rubrics as references, enabling effective RL on general reasoning tasks. Extensive experiments demonstrate the superiority of the proposed RuscaRL across various benchmarks, effectively expanding reasoning boundaries under the Best-of-N evaluation. Our code is available at https://github.com/IANNXANG/RuscaRL.
title Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.16949