Can Large Reasoning Models Self-Train?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shafayat, Sheikh, Tajwar, Fahim, Salakhutdinov, Ruslan, Schneider, Jeff, Zanette, Andrea
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912637143482368
author Shafayat, Sheikh
Tajwar, Fahim
Salakhutdinov, Ruslan
Schneider, Jeff
Zanette, Andrea
author_facet Shafayat, Sheikh
Tajwar, Fahim
Salakhutdinov, Ruslan
Schneider, Jeff
Zanette, Andrea
contents Recent successes of reinforcement learning (RL) in training large reasoning models motivate the question of whether self-training - the process where a model learns from its own judgments - can be sustained within RL. In this work, we study this question using majority voting as a simple self-feedback mechanism. On a comprehensive set of experiments on both synthetic and real reasoning tasks, we find that this basic approach improves not only the model's reasoning performance, but also its capability of generating better quality feedback for the next RL iteration, driving further model improvement. Yet our analysis also reveals a critical limitation of such a self-training paradigm - prolonged RL with self-reward leads to reward hacking where models learn to maximize training (pseudo-)reward, resulting in sudden and complete performance collapse. Together, these results highlight feedback design as the central challenge and call for future research on mechanisms to enable prolonged self-improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21444
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Large Reasoning Models Self-Train?
Shafayat, Sheikh
Tajwar, Fahim
Salakhutdinov, Ruslan
Schneider, Jeff
Zanette, Andrea
Machine Learning
Recent successes of reinforcement learning (RL) in training large reasoning models motivate the question of whether self-training - the process where a model learns from its own judgments - can be sustained within RL. In this work, we study this question using majority voting as a simple self-feedback mechanism. On a comprehensive set of experiments on both synthetic and real reasoning tasks, we find that this basic approach improves not only the model's reasoning performance, but also its capability of generating better quality feedback for the next RL iteration, driving further model improvement. Yet our analysis also reveals a critical limitation of such a self-training paradigm - prolonged RL with self-reward leads to reward hacking where models learn to maximize training (pseudo-)reward, resulting in sudden and complete performance collapse. Together, these results highlight feedback design as the central challenge and call for future research on mechanisms to enable prolonged self-improvement.
title Can Large Reasoning Models Self-Train?
topic Machine Learning
url https://arxiv.org/abs/2505.21444