Improving RL Exploration for LLM Reasoning through Retrospective Replay

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dou, Shihan, Wu, Muling, Xu, Jingwen, Zheng, Rui, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916825014468608
author Dou, Shihan
Wu, Muling
Xu, Jingwen
Zheng, Rui
Gui, Tao
Zhang, Qi
Huang, Xuanjing
author_facet Dou, Shihan
Wu, Muling
Xu, Jingwen
Zheng, Rui
Gui, Tao
Zhang, Qi
Huang, Xuanjing
contents Reinforcement learning (RL) has increasingly become a pivotal technique in the post-training of large language models (LLMs). The effective exploration of the output space is essential for the success of RL. We observe that for complex problems, during the early stages of training, the model exhibits strong exploratory capabilities and can identify promising solution ideas. However, its limited capability at this stage prevents it from successfully solving these problems. The early suppression of these potentially valuable solution ideas by the policy gradient hinders the model's ability to revisit and re-explore these ideas later. Consequently, although the LLM's capabilities improve in the later stages of training, it still struggles to effectively address these complex problems. To address this exploration issue, we propose a novel algorithm named Retrospective Replay-based Reinforcement Learning (RRL), which introduces a dynamic replay mechanism throughout the training process. RRL enables the model to revisit promising states identified in the early stages, thereby improving its efficiency and effectiveness in exploration. To evaluate the effectiveness of RRL, we conduct extensive experiments on complex reasoning tasks, including mathematical reasoning and code generation, and general dialogue tasks. The results indicate that RRL maintains high exploration efficiency throughout the training period, significantly enhancing the effectiveness of RL in optimizing LLMs for complicated reasoning tasks. Moreover, it also improves the performance of RLHF, making the model both safer and more helpful.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14363
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving RL Exploration for LLM Reasoning through Retrospective Replay
Dou, Shihan
Wu, Muling
Xu, Jingwen
Zheng, Rui
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Machine Learning
Computation and Language
Reinforcement learning (RL) has increasingly become a pivotal technique in the post-training of large language models (LLMs). The effective exploration of the output space is essential for the success of RL. We observe that for complex problems, during the early stages of training, the model exhibits strong exploratory capabilities and can identify promising solution ideas. However, its limited capability at this stage prevents it from successfully solving these problems. The early suppression of these potentially valuable solution ideas by the policy gradient hinders the model's ability to revisit and re-explore these ideas later. Consequently, although the LLM's capabilities improve in the later stages of training, it still struggles to effectively address these complex problems. To address this exploration issue, we propose a novel algorithm named Retrospective Replay-based Reinforcement Learning (RRL), which introduces a dynamic replay mechanism throughout the training process. RRL enables the model to revisit promising states identified in the early stages, thereby improving its efficiency and effectiveness in exploration. To evaluate the effectiveness of RRL, we conduct extensive experiments on complex reasoning tasks, including mathematical reasoning and code generation, and general dialogue tasks. The results indicate that RRL maintains high exploration efficiency throughout the training period, significantly enhancing the effectiveness of RL in optimizing LLMs for complicated reasoning tasks. Moreover, it also improves the performance of RLHF, making the model both safer and more helpful.
title Improving RL Exploration for LLM Reasoning through Retrospective Replay
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2504.14363