Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yang, Li, Sunzhu, Liu, Shunyu, Fang, Wenkai, Zhang, Kongcheng, Zhao, Jiale, Yang, Jingwen, Zhou, Yihe, Lv, Jianwei, Zheng, Tongya, Lu, Hengtong, Chen, Wei, Xie, Yan, Song, Mingli |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
by: Fang, Wenkai, et al.
Published: (2025)
by: Fang, Wenkai, et al.
Published: (2025)
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation
by: Li, Sunzhu, et al.
Published: (2026)
by: Li, Sunzhu, et al.
Published: (2026)
Odyssey: Empowering Minecraft Agents with Open-World Skills
by: Liu, Shunyu, et al.
Published: (2024)
by: Liu, Shunyu, et al.
Published: (2024)
Reasoning with Reinforced Functional Token Tuning
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?
by: Zhou, Yihe, et al.
Published: (2023)
by: Zhou, Yihe, et al.
Published: (2023)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Parallelized Planning-Acting for Efficient LLM-based Multi-Agent Systems in Minecraft
by: Li, Yaoru, et al.
Published: (2025)
by: Li, Yaoru, et al.
Published: (2025)
GraphScout: Empowering Large Language Models with Intrinsic Exploration Ability for Agentic Graph Reasoning
by: Ying, Yuchen, et al.
Published: (2026)
by: Ying, Yuchen, et al.
Published: (2026)
From GNNs to Trees: Multi-Granular Interpretability for Graph Neural Networks
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware Perspective
by: Qing, Yunpeng, et al.
Published: (2024)
by: Qing, Yunpeng, et al.
Published: (2024)
Interaction Pattern Disentangling for Multi-Agent Reinforcement Learning
by: Liu, Shunyu, et al.
Published: (2022)
by: Liu, Shunyu, et al.
Published: (2022)
Bi-level Mean Field: Dynamic Grouping for Large-Scale MARL
by: Zheng, Yuxuan, et al.
Published: (2025)
by: Zheng, Yuxuan, et al.
Published: (2025)
Unveiling Global Interactive Patterns across Graphs: Towards Interpretable Graph Neural Networks
by: Wang, Yuwen, et al.
Published: (2024)
by: Wang, Yuwen, et al.
Published: (2024)
COLA: Cross-city Mobility Transformer for Human Trajectory Simulation
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
A Survey of Direct Preference Optimization
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Temporal Prototype-Aware Learning for Active Voltage Control on Power Distribution Networks
by: Xu, Feiyang, et al.
Published: (2024)
by: Xu, Feiyang, et al.
Published: (2024)
Learning a Mini-batch Graph Transformer via Two-stage Interaction Augmentation
by: Li, Wenda, et al.
Published: (2024)
by: Li, Wenda, et al.
Published: (2024)
Simple Graph Condensation
by: Xiao, Zhenbang, et al.
Published: (2024)
by: Xiao, Zhenbang, et al.
Published: (2024)
ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
by: Li, Sunzhu, et al.
Published: (2025)
by: Li, Sunzhu, et al.
Published: (2025)
Curriculum Negative Mining For Temporal Networks
by: Chen, Ziyue, et al.
Published: (2024)
by: Chen, Ziyue, et al.
Published: (2024)
Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment
by: Lin, Zhiyu, et al.
Published: (2025)
by: Lin, Zhiyu, et al.
Published: (2025)
Reinforced Model Merging
by: Han, Jiaqi, et al.
Published: (2025)
by: Han, Jiaqi, et al.
Published: (2025)
A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges
by: Qing, Yunpeng, et al.
Published: (2022)
by: Qing, Yunpeng, et al.
Published: (2022)
Disentangled Condensation for Large-scale Graphs
by: Xiao, Zhenbang, et al.
Published: (2024)
by: Xiao, Zhenbang, et al.
Published: (2024)
Spatiotemporal-Augmented Graph Neural Networks for Human Mobility Simulation
by: Wang, Yu, et al.
Published: (2023)
by: Wang, Yu, et al.
Published: (2023)
On the Concept Trustworthiness in Concept Bottleneck Models
by: Huang, Qihan, et al.
Published: (2024)
by: Huang, Qihan, et al.
Published: (2024)
Holistic Semantic Representation for Navigational Trajectory Generation
by: Cao, Ji, et al.
Published: (2025)
by: Cao, Ji, et al.
Published: (2025)
Towards Efficient LLM-aware Heterogeneous Graph Learning
by: Li, Wenda, et al.
Published: (2025)
by: Li, Wenda, et al.
Published: (2025)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
Powerformer: A Section-adaptive Transformer for Power Flow Adjustment
by: Chen, Kaixuan, et al.
Published: (2024)
by: Chen, Kaixuan, et al.
Published: (2024)
Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs
by: Zhang, Wenjian, et al.
Published: (2026)
by: Zhang, Wenjian, et al.
Published: (2026)
Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric
by: Jia, Ruipeng, et al.
Published: (2026)
by: Jia, Ruipeng, et al.
Published: (2026)
Improving Adversarial Robustness via Feature Pattern Consistency Constraint
by: Hu, Jiacong, et al.
Published: (2024)
by: Hu, Jiacong, et al.
Published: (2024)
RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation
by: Cheng, Zelei, et al.
Published: (2024)
by: Cheng, Zelei, et al.
Published: (2024)
Step-wise Rubric Rewards for LLM Reasoning
by: Xie, Weichu, et al.
Published: (2026)
by: Xie, Weichu, et al.
Published: (2026)
The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental Design
by: Liu, Anjie, et al.
Published: (2026)
by: Liu, Anjie, et al.
Published: (2026)
DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
by: Wu, Yongtong, et al.
Published: (2026)
by: Wu, Yongtong, et al.
Published: (2026)
Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving
by: Song, Ziying, et al.
Published: (2025)
by: Song, Ziying, et al.
Published: (2025)
Similar Items
-
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
by: Fang, Wenkai, et al.
Published: (2025) -
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation
by: Li, Sunzhu, et al.
Published: (2026) -
Odyssey: Empowering Minecraft Agents with Open-World Skills
by: Liu, Shunyu, et al.
Published: (2024) -
Reasoning with Reinforced Functional Token Tuning
by: Zhang, Kongcheng, et al.
Published: (2025) -
Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?
by: Zhou, Yihe, et al.
Published: (2023)