You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roy, Shuvendu, Hajimirsadeghi, Hossein, Zhai, Mengyao, Samei, Golnoosh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917067336187904
author Roy, Shuvendu
Hajimirsadeghi, Hossein
Zhai, Mengyao
Samei, Golnoosh
author_facet Roy, Shuvendu
Hajimirsadeghi, Hossein
Zhai, Mengyao
Samei, Golnoosh
contents Recent advances in large language models have demonstrated the promise of unsupervised reinforcement learning (RL) methods for enhancing reasoning capabilities without external supervision. However, the generalizability of these label-free RL approaches to smaller base models with limited reasoning capabilities remains unexplored. In this work, we systematically investigate the performance of label-free RL methods across different model sizes and reasoning strengths, from 0.5B to 7B parameters. Our empirical analysis reveals critical limitations: label-free RL is highly dependent on the base model's pre-existing reasoning capability, with performance often degrading below baseline levels for weaker models. We find that smaller models fail to generate sufficiently long or diverse chain-of-thought reasoning to enable effective self-reflection, and that training data difficulty plays a crucial role in determining success. To address these challenges, we propose a simple yet effective method for label-free RL that utilizes curriculum learning to progressively introduce harder problems during training and mask no-majority rollouts during training. Additionally, we introduce a data curation pipeline to generate samples with predefined difficulty. Our approach demonstrates consistent improvements across all model sizes and reasoning capabilities, providing a path toward more robust unsupervised RL that can bootstrap reasoning abilities in resource-constrained models. We make our code available at https://github.com/BorealisAI/CuMa
format Preprint
id arxiv_https___arxiv_org_abs_2511_04902
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
Roy, Shuvendu
Hajimirsadeghi, Hossein
Zhai, Mengyao
Samei, Golnoosh
Machine Learning
Artificial Intelligence
Recent advances in large language models have demonstrated the promise of unsupervised reinforcement learning (RL) methods for enhancing reasoning capabilities without external supervision. However, the generalizability of these label-free RL approaches to smaller base models with limited reasoning capabilities remains unexplored. In this work, we systematically investigate the performance of label-free RL methods across different model sizes and reasoning strengths, from 0.5B to 7B parameters. Our empirical analysis reveals critical limitations: label-free RL is highly dependent on the base model's pre-existing reasoning capability, with performance often degrading below baseline levels for weaker models. We find that smaller models fail to generate sufficiently long or diverse chain-of-thought reasoning to enable effective self-reflection, and that training data difficulty plays a crucial role in determining success. To address these challenges, we propose a simple yet effective method for label-free RL that utilizes curriculum learning to progressively introduce harder problems during training and mask no-majority rollouts during training. Additionally, we introduce a data curation pipeline to generate samples with predefined difficulty. Our approach demonstrates consistent improvements across all model sizes and reasoning capabilities, providing a path toward more robust unsupervised RL that can bootstrap reasoning abilities in resource-constrained models. We make our code available at https://github.com/BorealisAI/CuMa
title You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.04902