Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ma, Lu, Liang, Hao, Qiang, Meiyi, Tang, Lexiang, Ma, Xiaochen, Wong, Zhen Hao, Niu, Junbo, Shen, Chengyu, He, Runming, Li, Yanhao, Cui, Bin, Zhang, Wentao
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912960807436288
author Ma, Lu
Liang, Hao
Qiang, Meiyi
Tang, Lexiang
Ma, Xiaochen
Wong, Zhen Hao
Niu, Junbo
Shen, Chengyu
He, Runming
Li, Yanhao
Cui, Bin
Zhang, Wentao
author_facet Ma, Lu
Liang, Hao
Qiang, Meiyi
Tang, Lexiang
Ma, Xiaochen
Wong, Zhen Hao
Niu, Junbo
Shen, Chengyu
He, Runming
Li, Yanhao
Cui, Bin
Zhang, Wentao
contents Recent advances in large language model (LLM) reasoning have shown that sophisticated behaviors such as planning and self-reflection can emerge through reinforcement learning (RL). However, despite these successes, RL in its current form remains insufficient to induce capabilities that exceed the limitations of the base model, as it is primarily optimized based on existing knowledge of the model rather than facilitating the acquisition of new information. To address this limitation, we employ supervised fine-tuning (SFT) to learn what RL cannot, which enables the incorporation of new knowledge and reasoning patterns by leveraging high-quality demonstration data. We analyze the training dynamics of RL and SFT for LLM reasoning and find that RL excels at maintaining and improving performance on questions within the model's original capabilities, while SFT is more effective at enabling progress on questions beyond the current scope of the model. Motivated by the complementary strengths of RL and SFT, we introduce a novel training approach, \textbf{ReLIFT} (\textbf{Re}inforcement \textbf{L}earning \textbf{I}nterleaved with Online \textbf{F}ine-\textbf{T}uning). In ReLIFT, the model is primarily trained using RL, but when it encounters challenging questions, high-quality solutions are collected for fine-tuning, and the training process alternates between RL and fine-tuning to enhance the model's reasoning abilities. ReLIFT achieves an average improvement of over +5.2 points across five competition-level benchmarks and one out-of-distribution benchmark compared to other zero-RL models. Furthermore, we demonstrate that ReLIFT outperforms both RL and SFT while using only 13\% of the detailed demonstration data, highlighting its scalability. These results provide compelling evidence that ReLIFT overcomes the fundamental limitations of RL and underscores the significant potential.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
Ma, Lu
Liang, Hao
Qiang, Meiyi
Tang, Lexiang
Ma, Xiaochen
Wong, Zhen Hao
Niu, Junbo
Shen, Chengyu
He, Runming
Li, Yanhao
Cui, Bin
Zhang, Wentao
Artificial Intelligence
Machine Learning
Recent advances in large language model (LLM) reasoning have shown that sophisticated behaviors such as planning and self-reflection can emerge through reinforcement learning (RL). However, despite these successes, RL in its current form remains insufficient to induce capabilities that exceed the limitations of the base model, as it is primarily optimized based on existing knowledge of the model rather than facilitating the acquisition of new information. To address this limitation, we employ supervised fine-tuning (SFT) to learn what RL cannot, which enables the incorporation of new knowledge and reasoning patterns by leveraging high-quality demonstration data. We analyze the training dynamics of RL and SFT for LLM reasoning and find that RL excels at maintaining and improving performance on questions within the model's original capabilities, while SFT is more effective at enabling progress on questions beyond the current scope of the model. Motivated by the complementary strengths of RL and SFT, we introduce a novel training approach, \textbf{ReLIFT} (\textbf{Re}inforcement \textbf{L}earning \textbf{I}nterleaved with Online \textbf{F}ine-\textbf{T}uning). In ReLIFT, the model is primarily trained using RL, but when it encounters challenging questions, high-quality solutions are collected for fine-tuning, and the training process alternates between RL and fine-tuning to enhance the model's reasoning abilities. ReLIFT achieves an average improvement of over +5.2 points across five competition-level benchmarks and one out-of-distribution benchmark compared to other zero-RL models. Furthermore, we demonstrate that ReLIFT outperforms both RL and SFT while using only 13\% of the detailed demonstration data, highlighting its scalability. These results provide compelling evidence that ReLIFT overcomes the fundamental limitations of RL and underscores the significant potential.
title Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.07527