Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hongru, Cai, Deng, Zhong, Wanjun, Huang, Shijue, Pan, Jeff Z., Liu, Zeming, Wong, Kam-Fai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908372051165184
author Wang, Hongru
Cai, Deng
Zhong, Wanjun
Huang, Shijue
Pan, Jeff Z.
Liu, Zeming
Wong, Kam-Fai
author_facet Wang, Hongru
Cai, Deng
Zhong, Wanjun
Huang, Shijue
Pan, Jeff Z.
Liu, Zeming
Wong, Kam-Fai
contents Inference-time scaling has attracted much attention which significantly enhance the performance of Large Language Models (LLMs) in complex reasoning tasks by increasing the length of Chain-of-Thought. These longer intermediate reasoning rationales embody various meta-reasoning skills in human cognition, such as reflection and decomposition, being difficult to create and acquire. In this work, we introduce \textit{Self-Reasoning Language Model} (SRLM), where the model itself can synthesize longer CoT data and iteratively improve performance through self-training. By incorporating a few demonstration examples (i.e., 1,000 samples) on how to unfold hidden reasoning chains from existing responses, which act as a reasoning catalyst, we demonstrate that SRLM not only enhances the model's initial performance but also ensures more stable and consistent improvements in subsequent iterations. Our proposed SRLM achieves an average absolute improvement of more than $+2.5$ points across five reasoning tasks: MMLU, GSM8K, ARC-C, HellaSwag, and BBH on two backbone models. Moreover, it brings more improvements with more times of sampling during inference, such as absolute $+7.89$ average improvement with $64$ sampling times, revealing the in-depth, diverse and creative reasoning paths in SRLM against the strong baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst
Wang, Hongru
Cai, Deng
Zhong, Wanjun
Huang, Shijue
Pan, Jeff Z.
Liu, Zeming
Wong, Kam-Fai
Computation and Language
Inference-time scaling has attracted much attention which significantly enhance the performance of Large Language Models (LLMs) in complex reasoning tasks by increasing the length of Chain-of-Thought. These longer intermediate reasoning rationales embody various meta-reasoning skills in human cognition, such as reflection and decomposition, being difficult to create and acquire. In this work, we introduce \textit{Self-Reasoning Language Model} (SRLM), where the model itself can synthesize longer CoT data and iteratively improve performance through self-training. By incorporating a few demonstration examples (i.e., 1,000 samples) on how to unfold hidden reasoning chains from existing responses, which act as a reasoning catalyst, we demonstrate that SRLM not only enhances the model's initial performance but also ensures more stable and consistent improvements in subsequent iterations. Our proposed SRLM achieves an average absolute improvement of more than $+2.5$ points across five reasoning tasks: MMLU, GSM8K, ARC-C, HellaSwag, and BBH on two backbone models. Moreover, it brings more improvements with more times of sampling during inference, such as absolute $+7.89$ average improvement with $64$ sampling times, revealing the in-depth, diverse and creative reasoning paths in SRLM against the strong baseline.
title Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst
topic Computation and Language
url https://arxiv.org/abs/2505.14116