Saved in:
Bibliographic Details
Main Authors: Wang, Ziliang, An, Kang, Zheng, Xuhui, Qian, Faqiang, Zhang, Weikun, Ouyang, Cijun, Cai, Jialu, Wang, Yuhang, Wu, Yichao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.00861
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908978974294016
author Wang, Ziliang
An, Kang
Zheng, Xuhui
Qian, Faqiang
Zhang, Weikun
Ouyang, Cijun
Cai, Jialu
Wang, Yuhang
Wu, Yichao
author_facet Wang, Ziliang
An, Kang
Zheng, Xuhui
Qian, Faqiang
Zhang, Weikun
Ouyang, Cijun
Cai, Jialu
Wang, Yuhang
Wu, Yichao
contents While search-augmented large language models (LLMs) exhibit impressive capabilities, their reliability in complex multi-hop reasoning remains limited. This limitation arises from three fundamental challenges: decomposition errors, where tasks are incorrectly broken down; retrieval missing, where key evidence fails to be retrieved; and reasoning errors, where flawed logic propagates through the reasoning chain. A single failure in any of these stages can derail the final answer. We propose Erasable Reinforcement Learning (ERL), a novel framework that transforms fragile reasoning into a robust process. ERL explicitly identifies faulty steps, erases them, and regenerates reasoning in place, preventing defective logic from propagating through the reasoning chain. This targeted correction mechanism turns brittle reasoning into a more resilient process. Models trained with ERL, termed ESearch, achieve substantial improvements on HotpotQA, MuSiQue, 2Wiki, and Bamboogle, with the 3B model achieving +8.48% EM and +11.56% F1, and the 7B model achieving +5.38% EM and +7.22% F1 over previous state-of-the-art(SOTA) results. These findings suggest that erasable reinforcement learning provides a powerful paradigm shift for robust multi-step reasoning in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
Wang, Ziliang
An, Kang
Zheng, Xuhui
Qian, Faqiang
Zhang, Weikun
Ouyang, Cijun
Cai, Jialu
Wang, Yuhang
Wu, Yichao
Computation and Language
Artificial Intelligence
Information Retrieval
While search-augmented large language models (LLMs) exhibit impressive capabilities, their reliability in complex multi-hop reasoning remains limited. This limitation arises from three fundamental challenges: decomposition errors, where tasks are incorrectly broken down; retrieval missing, where key evidence fails to be retrieved; and reasoning errors, where flawed logic propagates through the reasoning chain. A single failure in any of these stages can derail the final answer. We propose Erasable Reinforcement Learning (ERL), a novel framework that transforms fragile reasoning into a robust process. ERL explicitly identifies faulty steps, erases them, and regenerates reasoning in place, preventing defective logic from propagating through the reasoning chain. This targeted correction mechanism turns brittle reasoning into a more resilient process. Models trained with ERL, termed ESearch, achieve substantial improvements on HotpotQA, MuSiQue, 2Wiki, and Bamboogle, with the 3B model achieving +8.48% EM and +11.56% F1, and the 7B model achieving +5.38% EM and +7.22% F1 over previous state-of-the-art(SOTA) results. These findings suggest that erasable reinforcement learning provides a powerful paradigm shift for robust multi-step reasoning in LLMs.
title Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2510.00861