A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Mengqi, Zhao, Lei, So, Anthony Man-Cho, Sun, Ruoyu, Li, Xiao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914568302755840
author Li, Mengqi
Zhao, Lei
So, Anthony Man-Cho
Sun, Ruoyu
Li, Xiao
author_facet Li, Mengqi
Zhao, Lei
So, Anthony Man-Cho
Sun, Ruoyu
Li, Xiao
contents Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolving Post-Training (SePT), a simple post-training method that alternates between self-generation and training on self-generated responses. It repeatedly samples questions, uses the model itself to generate responses under a specified sampling temperature, and then trains the model on the self-generated data. In this self-training loop, we use an online data refresh mechanism, where each new batch is generated by the most recently updated model. Across six math reasoning benchmarks, SePT improves a strong no-training baseline, defined as the untuned base model evaluated at its best swept decoding temperature, on several tested models. Additional ablations demonstrate the importance of online data refresh and temperature dynamics. Overall, our results identify a practical regime where reasoning can be improved using self-generated supervision alone. Our code is available at https://github.com/ElementQi/SePT.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18814
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Li, Mengqi
Zhao, Lei
So, Anthony Man-Cho
Sun, Ruoyu
Li, Xiao
Machine Learning
Artificial Intelligence
Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolving Post-Training (SePT), a simple post-training method that alternates between self-generation and training on self-generated responses. It repeatedly samples questions, uses the model itself to generate responses under a specified sampling temperature, and then trains the model on the self-generated data. In this self-training loop, we use an online data refresh mechanism, where each new batch is generated by the most recently updated model. Across six math reasoning benchmarks, SePT improves a strong no-training baseline, defined as the untuned base model evaluated at its best swept decoding temperature, on several tested models. Additional ablations demonstrate the importance of online data refresh and temperature dynamics. Overall, our results identify a practical regime where reasoning can be improved using self-generated supervision alone. Our code is available at https://github.com/ElementQi/SePT.
title A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.18814