DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gu, Chenyang, Pu, Yewen, Yang, Bruce, Li, Xiaofan, Gao, Huan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911529977249792
author Gu, Chenyang
Pu, Yewen
Yang, Bruce
Li, Xiaofan
Gao, Huan
author_facet Gu, Chenyang
Pu, Yewen
Yang, Bruce
Li, Xiaofan
Gao, Huan
contents Enhancing LLMs with the ability to actively search external knowledge is crucial for complex and real-world tasks. Current approaches either rely on prompting to elicit the model's innate agent capabilities, or suffer from performance ceilings and collapse when applying RL to complex interactive tasks, leaving their true agentic potential untapped. To address this, we introduce \textbf{D}ynamic-filter \textbf{S}equence-level \textbf{P}olicy \textbf{O}ptimization (DSPO), an improved RL algorithm designed for robust agent training through sequence-level optimization and dynamic sample filtering. We train our model purely through RL to interleave multi-turn search and reasoning, obviating the need for supervised demonstration data. Across multiple QA benchmarks, our 7B model improves over a comparable previous work by \textbf{34.1\%}, and even outperforms the 14B model from previous work in complex multihop QA such as HotpotQA by nearly \textbf{9\% relative}, maintaining exceptional training stability.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09255
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
Gu, Chenyang
Pu, Yewen
Yang, Bruce
Li, Xiaofan
Gao, Huan
Computation and Language
Enhancing LLMs with the ability to actively search external knowledge is crucial for complex and real-world tasks. Current approaches either rely on prompting to elicit the model's innate agent capabilities, or suffer from performance ceilings and collapse when applying RL to complex interactive tasks, leaving their true agentic potential untapped. To address this, we introduce \textbf{D}ynamic-filter \textbf{S}equence-level \textbf{P}olicy \textbf{O}ptimization (DSPO), an improved RL algorithm designed for robust agent training through sequence-level optimization and dynamic sample filtering. We train our model purely through RL to interleave multi-turn search and reasoning, obviating the need for supervised demonstration data. Across multiple QA benchmarks, our 7B model improves over a comparable previous work by \textbf{34.1\%}, and even outperforms the 14B model from previous work in complex multihop QA such as HotpotQA by nearly \textbf{9\% relative}, maintaining exceptional training stability.
title DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
topic Computation and Language
url https://arxiv.org/abs/2510.09255