Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tang, Chenming, Huang, Hsiu-Yuan, Liu, Weijie, Bai, Clive, Yang, Saiyong, Wu, Yunfang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908968950956032
author Tang, Chenming
Huang, Hsiu-Yuan
Liu, Weijie
Bai, Clive
Yang, Saiyong
Wu, Yunfang
author_facet Tang, Chenming
Huang, Hsiu-Yuan
Liu, Weijie
Bai, Clive
Yang, Saiyong
Wu, Yunfang
contents Reinforcement learning with verifiable rewards (RLVR) has significantly boosted the reasoning capability of language models (LMs). However, existing RLVR approaches train LMs based on their own on-policy responses and are constrained by the initial capability of LMs, thus prone to exploration stagnation, in which LMs fail to solve more training problems and cannot further learn from the training data. Some approaches try to address this by leveraging off-policy solutions to training problems, but rely on external expert guidance that is limited in availability and scalability. In this work, we propose LTE (Learning to reason from Trial and Error), an approach that hints LMs with their previously self-made mistakes, not requiring any external expert guidance. Experiments validate the effectiveness of LTE, which outperforms the normal group relative policy optimization (GRPO) by 5.02 in Pass@1 and 9.96 in Pass@k on average across six mathematical reasoning benchmarks for Qwen3-8B-Base and even performs better than methods that require external guidance. Further analysis confirms that LTE successfully mitigates exploration stagnation and enhances both exploitation and exploration during training. Our code is available at https://github.com/JamyDon/LTE.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26109
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
Tang, Chenming
Huang, Hsiu-Yuan
Liu, Weijie
Bai, Clive
Yang, Saiyong
Wu, Yunfang
Machine Learning
Reinforcement learning with verifiable rewards (RLVR) has significantly boosted the reasoning capability of language models (LMs). However, existing RLVR approaches train LMs based on their own on-policy responses and are constrained by the initial capability of LMs, thus prone to exploration stagnation, in which LMs fail to solve more training problems and cannot further learn from the training data. Some approaches try to address this by leveraging off-policy solutions to training problems, but rely on external expert guidance that is limited in availability and scalability. In this work, we propose LTE (Learning to reason from Trial and Error), an approach that hints LMs with their previously self-made mistakes, not requiring any external expert guidance. Experiments validate the effectiveness of LTE, which outperforms the normal group relative policy optimization (GRPO) by 5.02 in Pass@1 and 9.96 in Pass@k on average across six mathematical reasoning benchmarks for Qwen3-8B-Base and even performs better than methods that require external guidance. Further analysis confirms that LTE successfully mitigates exploration stagnation and enhances both exploitation and exploration during training. Our code is available at https://github.com/JamyDon/LTE.
title Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
topic Machine Learning
url https://arxiv.org/abs/2510.26109