Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Junzhe, Xi, Zhiheng, Yang, Yajie, Luo, Hao, Dou, Shihan, Gui, Tao, Zhang, Qi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914491035287552
author Wang, Junzhe
Xi, Zhiheng
Yang, Yajie
Luo, Hao
Dou, Shihan
Gui, Tao
Zhang, Qi
author_facet Wang, Junzhe
Xi, Zhiheng
Yang, Yajie
Luo, Hao
Dou, Shihan
Gui, Tao
Zhang, Qi
contents Search agents extend Large Language Models (LLMs) beyond static parametric knowledge by enabling access to up-to-date and long-tail information unavailable during pretraining. While reinforcement learning has been widely adopted for training such agents, existing approaches face key limitations: process supervision often suffers from unstable value estimation, whereas outcome supervision struggles with credit assignment due to sparse, trajectory-level rewards. To bridge this gap, we propose Contribution-Weighted GRPO (CW-GRPO), a framework that integrates process supervision into group relative policy optimization. Instead of directly optimizing process rewards, CW-GRPO employs an LLM judge to assess the retrieval utility and reasoning correctness at each search round, producing per-round contribution scores. These scores are used to rescale outcome-based advantages along the trajectory, enabling fine-grained credit assignment without sacrificing optimization stability. Experiments on multiple knowledge-intensive benchmarks show that CW-GRPO outperforms standard GRPO by 5.0% on Qwen3-8B and 6.3% on Qwen3-1.7B, leading to more effective search behaviors. Additional analysis reveals that successful trajectories exhibit concentrated contributions in specific rounds, providing empirical insight into search agent tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14267
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
Wang, Junzhe
Xi, Zhiheng
Yang, Yajie
Luo, Hao
Dou, Shihan
Gui, Tao
Zhang, Qi
Machine Learning
Artificial Intelligence
Search agents extend Large Language Models (LLMs) beyond static parametric knowledge by enabling access to up-to-date and long-tail information unavailable during pretraining. While reinforcement learning has been widely adopted for training such agents, existing approaches face key limitations: process supervision often suffers from unstable value estimation, whereas outcome supervision struggles with credit assignment due to sparse, trajectory-level rewards. To bridge this gap, we propose Contribution-Weighted GRPO (CW-GRPO), a framework that integrates process supervision into group relative policy optimization. Instead of directly optimizing process rewards, CW-GRPO employs an LLM judge to assess the retrieval utility and reasoning correctness at each search round, producing per-round contribution scores. These scores are used to rescale outcome-based advantages along the trajectory, enabling fine-grained credit assignment without sacrificing optimization stability. Experiments on multiple knowledge-intensive benchmarks show that CW-GRPO outperforms standard GRPO by 5.0% on Qwen3-8B and 6.3% on Qwen3-1.7B, leading to more effective search behaviors. Additional analysis reveals that successful trajectories exhibit concentrated contributions in specific rounds, providing empirical insight into search agent tasks.
title Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.14267