Repurposing Synthetic Data for Fine-grained Search Agent Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yida, Li, Kuan, Wu, Xixi, Zhang, Liwen, Zhang, Dingchu, Li, Baixuan, Song, Maojia, Chen, Zhuo, Wang, Chenxi, Wang, Xinyu, Tu, Kewei, Xie, Pengjun, Zhou, Jingren, Jiang, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910030801928192
author Zhao, Yida
Li, Kuan
Wu, Xixi
Zhang, Liwen
Zhang, Dingchu
Li, Baixuan
Song, Maojia
Chen, Zhuo
Wang, Chenxi
Wang, Xinyu
Tu, Kewei
Xie, Pengjun
Zhou, Jingren
Jiang, Yong
author_facet Zhao, Yida
Li, Kuan
Wu, Xixi
Zhang, Liwen
Zhang, Dingchu
Li, Baixuan
Song, Maojia
Chen, Zhuo
Wang, Chenxi
Wang, Xinyu
Tu, Kewei
Xie, Pengjun
Zhou, Jingren
Jiang, Yong
contents LLM-based search agents are increasingly trained on entity-centric synthetic data to solve complex, knowledge-intensive tasks. However, prevailing training methods like Group Relative Policy Optimization (GRPO) discard this rich entity information, relying instead on sparse, outcome-based rewards. This critical limitation renders them unable to distinguish informative "near-miss" samples-those with substantially correct reasoning but a flawed final answer-from complete failures, thus discarding valuable learning signals. We address this by leveraging the very entities discarded during training. Our empirical analysis reveals a strong positive correlation between the number of ground-truth entities identified during an agent's reasoning process and final answer accuracy. Building on this insight, we introduce Entity-aware Group Relative Policy Optimization (E-GRPO), a novel framework that formulates a dense entity-aware reward function. E-GRPO assigns partial rewards to incorrect samples proportional to their entity match rate, enabling the model to effectively learn from these "near-misses". Experiments on diverse question-answering (QA) and deep research benchmarks show that E-GRPO consistently and significantly outperforms the GRPO baseline. Furthermore, our analysis reveals that E-GRPO not only achieves superior accuracy but also induces more efficient reasoning policies that require fewer tool calls, demonstrating a more effective and sample-efficient approach to aligning search agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Repurposing Synthetic Data for Fine-grained Search Agent Supervision
Zhao, Yida
Li, Kuan
Wu, Xixi
Zhang, Liwen
Zhang, Dingchu
Li, Baixuan
Song, Maojia
Chen, Zhuo
Wang, Chenxi
Wang, Xinyu
Tu, Kewei
Xie, Pengjun
Zhou, Jingren
Jiang, Yong
Computation and Language
Artificial Intelligence
LLM-based search agents are increasingly trained on entity-centric synthetic data to solve complex, knowledge-intensive tasks. However, prevailing training methods like Group Relative Policy Optimization (GRPO) discard this rich entity information, relying instead on sparse, outcome-based rewards. This critical limitation renders them unable to distinguish informative "near-miss" samples-those with substantially correct reasoning but a flawed final answer-from complete failures, thus discarding valuable learning signals. We address this by leveraging the very entities discarded during training. Our empirical analysis reveals a strong positive correlation between the number of ground-truth entities identified during an agent's reasoning process and final answer accuracy. Building on this insight, we introduce Entity-aware Group Relative Policy Optimization (E-GRPO), a novel framework that formulates a dense entity-aware reward function. E-GRPO assigns partial rewards to incorrect samples proportional to their entity match rate, enabling the model to effectively learn from these "near-misses". Experiments on diverse question-answering (QA) and deep research benchmarks show that E-GRPO consistently and significantly outperforms the GRPO baseline. Furthermore, our analysis reveals that E-GRPO not only achieves superior accuracy but also induces more efficient reasoning policies that require fewer tool calls, demonstrating a more effective and sample-efficient approach to aligning search agents.
title Repurposing Synthetic Data for Fine-grained Search Agent Supervision
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.24694