CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Hansi, Collins, Liam, Kumar, Bhuvesh, Shah, Neil, Zamani, Hamed
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911613781540864
author Zeng, Hansi
Collins, Liam
Kumar, Bhuvesh
Shah, Neil
Zamani, Hamed
author_facet Zeng, Hansi
Collins, Liam
Kumar, Bhuvesh
Shah, Neil
Zamani, Hamed
contents Agentic search -- the task of training agents that iteratively reason, issue queries, and synthesize retrieved information to answer complex questions -- has achieved remarkable progress through reinforcement learning (RL). However, existing approaches such as Search-R1, treat the retrieval system as a fixed tool, optimizing only the reasoning agent while the retrieval component remains unchanged. A preliminary experiment reveals that the gap between an oracle and a fixed retrieval system reaches up to +26.8% relative F1 improvement across seven QA benchmarks, suggesting that the retrieval system is a key bottleneck in scaling agentic search performance. Motivated by this finding, we propose CoSearch, a framework that jointly trains a multi-step reasoning agent and a generative document ranking model via Group Relative Policy Optimization (GRPO). To enable effective GRPO training for the ranker -- whose inputs vary across reasoning trajectories -- we introduce a semantic grouping strategy that clusters sub-queries by token-level similarity, forming valid optimization groups without additional rollouts. We further design a composite reward combining ranking quality signals with trajectory-level outcome feedback, providing the ranker with both immediate and long-term learning signals. Experiments on seven single-hop and multi-hop QA benchmarks demonstrate consistent improvements over strong baselines, with ablation studies validating each design choice. Our results show that joint training of the reasoning agent and retrieval system is both feasible and strongly performant, pointing to a key ingredient for future search agents.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17555
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search
Zeng, Hansi
Collins, Liam
Kumar, Bhuvesh
Shah, Neil
Zamani, Hamed
Artificial Intelligence
Computation and Language
Information Retrieval
Agentic search -- the task of training agents that iteratively reason, issue queries, and synthesize retrieved information to answer complex questions -- has achieved remarkable progress through reinforcement learning (RL). However, existing approaches such as Search-R1, treat the retrieval system as a fixed tool, optimizing only the reasoning agent while the retrieval component remains unchanged. A preliminary experiment reveals that the gap between an oracle and a fixed retrieval system reaches up to +26.8% relative F1 improvement across seven QA benchmarks, suggesting that the retrieval system is a key bottleneck in scaling agentic search performance. Motivated by this finding, we propose CoSearch, a framework that jointly trains a multi-step reasoning agent and a generative document ranking model via Group Relative Policy Optimization (GRPO). To enable effective GRPO training for the ranker -- whose inputs vary across reasoning trajectories -- we introduce a semantic grouping strategy that clusters sub-queries by token-level similarity, forming valid optimization groups without additional rollouts. We further design a composite reward combining ranking quality signals with trajectory-level outcome feedback, providing the ranker with both immediate and long-term learning signals. Experiments on seven single-hop and multi-hop QA benchmarks demonstrate consistent improvements over strong baselines, with ablation studies validating each design choice. Our results show that joint training of the reasoning agent and retrieval system is both feasible and strongly performant, pointing to a key ingredient for future search agents.
title CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search
topic Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2604.17555