Effective Learning for Small Reasoning Models: An Empirical Study on 0.5B Reasoning LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Xialie, Ma, Peixian, Jia, Zhikai, Cao, Zane, Liu, Shiwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912715824431104
author Zhuang, Xialie
Ma, Peixian
Jia, Zhikai
Cao, Zane
Liu, Shiwei
author_facet Zhuang, Xialie
Ma, Peixian
Jia, Zhikai
Cao, Zane
Liu, Shiwei
contents The ongoing evolution of language models has led to the development of large-scale architectures that demonstrate exceptional performance across a wide range of tasks. However, these models come with significant computational and energy demands, as well as potential privacy implications. In this context, Small Reasoning Language Models (SRLMs) with approximately 0.5 billion parameters present a compelling alternative due to their remarkable computational efficiency and cost-effectiveness, particularly in resource-constrained environments. Despite these advantages, the limited capacity of 0.5 billion parameter models poses challenges in handling complex tasks such as mathematical reasoning. This research investigates various training strategies, including supervised fine-tuning (SFT), knowledge distillation (KD), and reinforcement learning (RL), as well as their hybrid implementations, to enhance the performance of 0.5B SRLMs. We analyze effective methodologies to bridge the performance gap between SRLMS and larger models and present insights into optimal training pipelines tailored for these smaller architectures. Through extensive experimental validation and analysis, our work aims to provide actionable recommendations for maximizing the reasoning capabilities of 0.5B models.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13404
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Effective Learning for Small Reasoning Models: An Empirical Study on 0.5B Reasoning LLMs
Zhuang, Xialie
Ma, Peixian
Jia, Zhikai
Cao, Zane
Liu, Shiwei
Artificial Intelligence
The ongoing evolution of language models has led to the development of large-scale architectures that demonstrate exceptional performance across a wide range of tasks. However, these models come with significant computational and energy demands, as well as potential privacy implications. In this context, Small Reasoning Language Models (SRLMs) with approximately 0.5 billion parameters present a compelling alternative due to their remarkable computational efficiency and cost-effectiveness, particularly in resource-constrained environments. Despite these advantages, the limited capacity of 0.5 billion parameter models poses challenges in handling complex tasks such as mathematical reasoning. This research investigates various training strategies, including supervised fine-tuning (SFT), knowledge distillation (KD), and reinforcement learning (RL), as well as their hybrid implementations, to enhance the performance of 0.5B SRLMs. We analyze effective methodologies to bridge the performance gap between SRLMS and larger models and present insights into optimal training pipelines tailored for these smaller architectures. Through extensive experimental validation and analysis, our work aims to provide actionable recommendations for maximizing the reasoning capabilities of 0.5B models.
title Effective Learning for Small Reasoning Models: An Empirical Study on 0.5B Reasoning LLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2506.13404