CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yaocheng, Huang, Haohuan, Song, Zijun, Zhu, Yuanheng, Zhang, Qichao, Zhao, Zijie, Zhao, Dongbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914159867723776
author Zhang, Yaocheng
Huang, Haohuan
Song, Zijun
Zhu, Yuanheng
Zhang, Qichao
Zhao, Zijie
Zhao, Dongbin
author_facet Zhang, Yaocheng
Huang, Haohuan
Song, Zijun
Zhu, Yuanheng
Zhang, Qichao
Zhao, Zijie
Zhao, Dongbin
contents Tool-Integrated Reasoning (TIR) with search engines enables large language models to iteratively retrieve up-to-date external knowledge, enhancing adaptability and generalization in complex question-answering tasks. However, existing search agent pipelines typically depend on reinforcement learning based optimization, which often suffers from sparse outcome rewards, leading to inefficient exploration and unstable training. We introduce CriticSearch, a fine-grained credit-assignment framework that supplies dense, turn-level feedback via a retrospective critic mechanism. During training, a frozen, asymmetric critique LLM retrospectively evaluates each turn using privileged information from the full trajectory and gold answers, converting these assessments into stable, dense rewards that guide policy improvement. Experimental results across diverse multi-hop reasoning benchmarks demonstrate that CriticSearch consistently outperforms existing baselines, achieving faster convergence, improved training stability, and higher performance.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
Zhang, Yaocheng
Huang, Haohuan
Song, Zijun
Zhu, Yuanheng
Zhang, Qichao
Zhao, Zijie
Zhao, Dongbin
Computation and Language
Tool-Integrated Reasoning (TIR) with search engines enables large language models to iteratively retrieve up-to-date external knowledge, enhancing adaptability and generalization in complex question-answering tasks. However, existing search agent pipelines typically depend on reinforcement learning based optimization, which often suffers from sparse outcome rewards, leading to inefficient exploration and unstable training. We introduce CriticSearch, a fine-grained credit-assignment framework that supplies dense, turn-level feedback via a retrospective critic mechanism. During training, a frozen, asymmetric critique LLM retrospectively evaluates each turn using privileged information from the full trajectory and gold answers, converting these assessments into stable, dense rewards that guide policy improvement. Experimental results across diverse multi-hop reasoning benchmarks demonstrate that CriticSearch consistently outperforms existing baselines, achieving faster convergence, improved training stability, and higher performance.
title CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
topic Computation and Language
url https://arxiv.org/abs/2511.12159