Repurposing Synthetic Data for Fine-grained Search Agent Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910030801928192 |
|---|---|
| author | Zhao, Yida Li, Kuan Wu, Xixi Zhang, Liwen Zhang, Dingchu Li, Baixuan Song, Maojia Chen, Zhuo Wang, Chenxi Wang, Xinyu Tu, Kewei Xie, Pengjun Zhou, Jingren Jiang, Yong |
| author_facet | Zhao, Yida Li, Kuan Wu, Xixi Zhang, Liwen Zhang, Dingchu Li, Baixuan Song, Maojia Chen, Zhuo Wang, Chenxi Wang, Xinyu Tu, Kewei Xie, Pengjun Zhou, Jingren Jiang, Yong |
| contents | LLM-based search agents are increasingly trained on entity-centric synthetic data to solve complex, knowledge-intensive tasks. However, prevailing training methods like Group Relative Policy Optimization (GRPO) discard this rich entity information, relying instead on sparse, outcome-based rewards. This critical limitation renders them unable to distinguish informative "near-miss" samples-those with substantially correct reasoning but a flawed final answer-from complete failures, thus discarding valuable learning signals. We address this by leveraging the very entities discarded during training. Our empirical analysis reveals a strong positive correlation between the number of ground-truth entities identified during an agent's reasoning process and final answer accuracy. Building on this insight, we introduce Entity-aware Group Relative Policy Optimization (E-GRPO), a novel framework that formulates a dense entity-aware reward function. E-GRPO assigns partial rewards to incorrect samples proportional to their entity match rate, enabling the model to effectively learn from these "near-misses". Experiments on diverse question-answering (QA) and deep research benchmarks show that E-GRPO consistently and significantly outperforms the GRPO baseline. Furthermore, our analysis reveals that E-GRPO not only achieves superior accuracy but also induces more efficient reasoning policies that require fewer tool calls, demonstrating a more effective and sample-efficient approach to aligning search agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_24694 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Repurposing Synthetic Data for Fine-grained Search Agent Supervision Zhao, Yida Li, Kuan Wu, Xixi Zhang, Liwen Zhang, Dingchu Li, Baixuan Song, Maojia Chen, Zhuo Wang, Chenxi Wang, Xinyu Tu, Kewei Xie, Pengjun Zhou, Jingren Jiang, Yong Computation and Language Artificial Intelligence LLM-based search agents are increasingly trained on entity-centric synthetic data to solve complex, knowledge-intensive tasks. However, prevailing training methods like Group Relative Policy Optimization (GRPO) discard this rich entity information, relying instead on sparse, outcome-based rewards. This critical limitation renders them unable to distinguish informative "near-miss" samples-those with substantially correct reasoning but a flawed final answer-from complete failures, thus discarding valuable learning signals. We address this by leveraging the very entities discarded during training. Our empirical analysis reveals a strong positive correlation between the number of ground-truth entities identified during an agent's reasoning process and final answer accuracy. Building on this insight, we introduce Entity-aware Group Relative Policy Optimization (E-GRPO), a novel framework that formulates a dense entity-aware reward function. E-GRPO assigns partial rewards to incorrect samples proportional to their entity match rate, enabling the model to effectively learn from these "near-misses". Experiments on diverse question-answering (QA) and deep research benchmarks show that E-GRPO consistently and significantly outperforms the GRPO baseline. Furthermore, our analysis reveals that E-GRPO not only achieves superior accuracy but also induces more efficient reasoning policies that require fewer tool calls, demonstrating a more effective and sample-efficient approach to aligning search agents. |
| title | Repurposing Synthetic Data for Fine-grained Search Agent Supervision |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.24694 |