ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xixi, Li, Kuan, Zhao, Yida, Zhang, Liwen, Ou, Litu, Yin, Huifeng, Zhang, Zhongwang, Yu, Xinmiao, Zhang, Dingchu, Jiang, Yong, Xie, Pengjun, Huang, Fei, Cheng, Minhao, Wang, Shuai, Cheng, Hong, Zhou, Jingren
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908913602920448
author Wu, Xixi
Li, Kuan
Zhao, Yida
Zhang, Liwen
Ou, Litu
Yin, Huifeng
Zhang, Zhongwang
Yu, Xinmiao
Zhang, Dingchu
Jiang, Yong
Xie, Pengjun
Huang, Fei
Cheng, Minhao
Wang, Shuai
Cheng, Hong
Zhou, Jingren
author_facet Wu, Xixi
Li, Kuan
Zhao, Yida
Zhang, Liwen
Ou, Litu
Yin, Huifeng
Zhang, Zhongwang
Yu, Xinmiao
Zhang, Dingchu
Jiang, Yong
Xie, Pengjun
Huang, Fei
Cheng, Minhao
Wang, Shuai
Cheng, Hong
Zhou, Jingren
contents Large Language Model (LLM)-based web agents excel at knowledge-intensive tasks but face a fundamental conflict between the need for extensive exploration and the constraints of limited context windows. Current solutions typically rely on architectural modifications, e.g., internal memory tokens, which break compatibility with pre-existing agents and necessitate costly end-to-end retraining. To overcome these limitations, we introduce ReSum, a lightweight, plug-and-play paradigm that enables unbounded exploration by periodically invoking an external tool to condense interaction histories into compact summaries. Although this paradigm functions without training, standard agents are not inherently aligned to reason over such compressed contexts. To bridge this gap, we propose ReSum-GRPO, which adapts Group Relative Policy Optimization (GRPO) via advantage broadcasting to propagate final rewards across segmented trajectories, enabling credit assignments over long-horizons. Extensive experiments show that ReSum achieves a 4.5% improvement over ReAct in training-free settings, with ReSum-GRPO yielding a further 8.2% gain. Notably, with only 1K training samples, a ReSum-enhanced 30B agent achieves competitive performance with leading open-source models, showing ReSum's effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13313
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
Wu, Xixi
Li, Kuan
Zhao, Yida
Zhang, Liwen
Ou, Litu
Yin, Huifeng
Zhang, Zhongwang
Yu, Xinmiao
Zhang, Dingchu
Jiang, Yong
Xie, Pengjun
Huang, Fei
Cheng, Minhao
Wang, Shuai
Cheng, Hong
Zhou, Jingren
Computation and Language
Large Language Model (LLM)-based web agents excel at knowledge-intensive tasks but face a fundamental conflict between the need for extensive exploration and the constraints of limited context windows. Current solutions typically rely on architectural modifications, e.g., internal memory tokens, which break compatibility with pre-existing agents and necessitate costly end-to-end retraining. To overcome these limitations, we introduce ReSum, a lightweight, plug-and-play paradigm that enables unbounded exploration by periodically invoking an external tool to condense interaction histories into compact summaries. Although this paradigm functions without training, standard agents are not inherently aligned to reason over such compressed contexts. To bridge this gap, we propose ReSum-GRPO, which adapts Group Relative Policy Optimization (GRPO) via advantage broadcasting to propagate final rewards across segmented trajectories, enabling credit assignments over long-horizons. Extensive experiments show that ReSum achieves a 4.5% improvement over ReAct in training-free settings, with ReSum-GRPO yielding a further 8.2% gain. Notably, with only 1K training samples, a ReSum-enhanced 30B agent achieves competitive performance with leading open-source models, showing ReSum's effectiveness.
title ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
topic Computation and Language
url https://arxiv.org/abs/2509.13313