Saved in:
Bibliographic Details
Main Authors: Ji, Yunjie, Zhao, Sitong, Tian, Xiaoyu, Wang, Haotian, Chen, Shuaiting, Peng, Yiping, Zhao, Han, Li, Xiangang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.00829
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916669780131840
author Ji, Yunjie
Zhao, Sitong
Tian, Xiaoyu
Wang, Haotian
Chen, Shuaiting
Peng, Yiping
Zhao, Han
Li, Xiangang
author_facet Ji, Yunjie
Zhao, Sitong
Tian, Xiaoyu
Wang, Haotian
Chen, Shuaiting
Peng, Yiping
Zhao, Han
Li, Xiangang
contents Enhancing the reasoning capabilities of Large Language Models (LLMs) with efficiency and scalability remains a fundamental challenge in artificial intelligence research. This paper presents a rigorous experimental investigation into how difficulty-aware staged reinforcement learning (RL) strategies can substantially improve LLM reasoning performance. Through systematic analysis, we demonstrate that strategically selecting training data according to well-defined difficulty levels markedly enhances RL optimization. Moreover, we introduce a staged training methodology, progressively exposing models to increasingly challenging tasks, further amplifying reasoning capabilities. Our findings reveal significant cross-domain benefits when simultaneously training models on mathematical reasoning and code generation tasks. Notably, our proposed approach enables a 1.5B parameter model to achieve an accuracy of 42.3\% on the AIME-2024 benchmark, 89.5\% on the MATH-500 benchmark. These results underscore the efficacy of our method in advancing the reasoning proficiency of LLMs. We will open-source our datasets on GitHub and Hugging Face.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00829
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
Ji, Yunjie
Zhao, Sitong
Tian, Xiaoyu
Wang, Haotian
Chen, Shuaiting
Peng, Yiping
Zhao, Han
Li, Xiangang
Computation and Language
Enhancing the reasoning capabilities of Large Language Models (LLMs) with efficiency and scalability remains a fundamental challenge in artificial intelligence research. This paper presents a rigorous experimental investigation into how difficulty-aware staged reinforcement learning (RL) strategies can substantially improve LLM reasoning performance. Through systematic analysis, we demonstrate that strategically selecting training data according to well-defined difficulty levels markedly enhances RL optimization. Moreover, we introduce a staged training methodology, progressively exposing models to increasingly challenging tasks, further amplifying reasoning capabilities. Our findings reveal significant cross-domain benefits when simultaneously training models on mathematical reasoning and code generation tasks. Notably, our proposed approach enables a 1.5B parameter model to achieve an accuracy of 42.3\% on the AIME-2024 benchmark, 89.5\% on the MATH-500 benchmark. These results underscore the efficacy of our method in advancing the reasoning proficiency of LLMs. We will open-source our datasets on GitHub and Hugging Face.
title How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
topic Computation and Language
url https://arxiv.org/abs/2504.00829