Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Parashar, Shubham, Gui, Shurui, Li, Xiner, Ling, Hongyi, Vemuri, Sushil, Olson, Blake, Li, Eric, Zhang, Yu, Caverlee, James, Kalathil, Dileep, Ji, Shuiwang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917346223849472
author Parashar, Shubham
Gui, Shurui
Li, Xiner
Ling, Hongyi
Vemuri, Sushil
Olson, Blake
Li, Eric
Zhang, Yu
Caverlee, James
Kalathil, Dileep
Ji, Shuiwang
author_facet Parashar, Shubham
Gui, Shurui
Li, Xiner
Ling, Hongyi
Vemuri, Sushil
Olson, Blake
Li, Eric
Zhang, Yu
Caverlee, James
Kalathil, Dileep
Ji, Shuiwang
contents We aim to improve the reasoning capabilities of language models via reinforcement learning (RL). Recent RL post-trained models like DeepSeek-R1 have demonstrated reasoning abilities on mathematical and coding tasks. However, prior studies suggest that using RL alone to improve reasoning on inherently difficult tasks is less effective. Here, we draw inspiration from curriculum learning and propose to schedule tasks from easy to hard (E2H), allowing LLMs to build reasoning skills gradually. Our method is termed E2H Reasoner. Empirically, we observe that, although easy tasks are important initially, fading them out through appropriate scheduling is essential in preventing overfitting. Theoretically, we establish convergence guarantees for E2H Reasoner within an approximate policy iteration framework. We derive finite-sample complexity bounds and show that when tasks are appropriately decomposed and conditioned, learning through curriculum stages requires fewer total samples than direct learning. Experiments across multiple domains show that E2H Reasoner significantly improves the reasoning ability of small LLMs (1.5B to 3B), which otherwise struggle when trained with vanilla RL alone, highlighting the effectiveness of our method. Our code can be found on https://github.com/divelab/E2H-Reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
Parashar, Shubham
Gui, Shurui
Li, Xiner
Ling, Hongyi
Vemuri, Sushil
Olson, Blake
Li, Eric
Zhang, Yu
Caverlee, James
Kalathil, Dileep
Ji, Shuiwang
Machine Learning
Artificial Intelligence
Computation and Language
We aim to improve the reasoning capabilities of language models via reinforcement learning (RL). Recent RL post-trained models like DeepSeek-R1 have demonstrated reasoning abilities on mathematical and coding tasks. However, prior studies suggest that using RL alone to improve reasoning on inherently difficult tasks is less effective. Here, we draw inspiration from curriculum learning and propose to schedule tasks from easy to hard (E2H), allowing LLMs to build reasoning skills gradually. Our method is termed E2H Reasoner. Empirically, we observe that, although easy tasks are important initially, fading them out through appropriate scheduling is essential in preventing overfitting. Theoretically, we establish convergence guarantees for E2H Reasoner within an approximate policy iteration framework. We derive finite-sample complexity bounds and show that when tasks are appropriately decomposed and conditioned, learning through curriculum stages requires fewer total samples than direct learning. Experiments across multiple domains show that E2H Reasoner significantly improves the reasoning ability of small LLMs (1.5B to 3B), which otherwise struggle when trained with vanilla RL alone, highlighting the effectiveness of our method. Our code can be found on https://github.com/divelab/E2H-Reasoning.
title Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.06632