Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Xiao-Wen, Zhu, Xuan-Yi, Wei, Wen-Da, Zhang, Ding-Chu, Shao, Jie-Jing, Zhou, Zhi, Guo, Lan-Zhe, Li, Yu-Feng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917915311210496
author Yang, Xiao-Wen
Zhu, Xuan-Yi
Wei, Wen-Da
Zhang, Ding-Chu
Shao, Jie-Jing
Zhou, Zhi
Guo, Lan-Zhe
Li, Yu-Feng
author_facet Yang, Xiao-Wen
Zhu, Xuan-Yi
Wei, Wen-Da
Zhang, Ding-Chu
Shao, Jie-Jing
Zhou, Zhi
Guo, Lan-Zhe
Li, Yu-Feng
contents The integration of slow-thinking mechanisms into large language models (LLMs) offers a promising way toward achieving Level 2 AGI Reasoners, as exemplified by systems like OpenAI's o1. However, several significant challenges remain, including inefficient overthinking and an overreliance on auxiliary reward models. We point out that these limitations stem from LLMs' inability to internalize the search process, a key component of effective reasoning. A critical step toward addressing this issue is enabling LLMs to autonomously determine when and where to backtrack, a fundamental operation in traditional search algorithms. To this end, we propose a self-backtracking mechanism that equips LLMs with the ability to backtrack during both training and inference. This mechanism not only enhances reasoning ability but also efficiency by transforming slow-thinking processes into fast-thinking through self-improvement. Empirical evaluations demonstrate that our proposal significantly enhances the reasoning capabilities of LLMs, achieving a performance gain of over 40 percent compared to the optimal-path supervised fine-tuning method. We believe this study introduces a novel and promising pathway for developing more advanced and robust Reasoners.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04404
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
Yang, Xiao-Wen
Zhu, Xuan-Yi
Wei, Wen-Da
Zhang, Ding-Chu
Shao, Jie-Jing
Zhou, Zhi
Guo, Lan-Zhe
Li, Yu-Feng
Computation and Language
Artificial Intelligence
The integration of slow-thinking mechanisms into large language models (LLMs) offers a promising way toward achieving Level 2 AGI Reasoners, as exemplified by systems like OpenAI's o1. However, several significant challenges remain, including inefficient overthinking and an overreliance on auxiliary reward models. We point out that these limitations stem from LLMs' inability to internalize the search process, a key component of effective reasoning. A critical step toward addressing this issue is enabling LLMs to autonomously determine when and where to backtrack, a fundamental operation in traditional search algorithms. To this end, we propose a self-backtracking mechanism that equips LLMs with the ability to backtrack during both training and inference. This mechanism not only enhances reasoning ability but also efficiency by transforming slow-thinking processes into fast-thinking through self-improvement. Empirical evaluations demonstrate that our proposal significantly enhances the reasoning capabilities of LLMs, achieving a performance gain of over 40 percent compared to the optimal-path supervised fine-tuning method. We believe this study introduces a novel and promising pathway for developing more advanced and robust Reasoners.
title Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.04404