InfoFlow: Reinforcing Search Agent Via Reward Density Optimization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Kun, Qian, Hongjin, Liu, Zheng, Xia, Ziyi, Xiao, Shitao, Bao, Siqi, Zhao, Jun, Liu, Kang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917050466697216
author Luo, Kun
Qian, Hongjin
Liu, Zheng
Xia, Ziyi
Xiao, Shitao
Bao, Siqi
Zhao, Jun
Liu, Kang
author_facet Luo, Kun
Qian, Hongjin
Liu, Zheng
Xia, Ziyi
Xiao, Shitao
Bao, Siqi
Zhao, Jun
Liu, Kang
contents Reinforcement Learning with Verifiable Rewards (RLVR) is a promising approach for enhancing agentic deep search. However, its application is often hindered by low \textbf{Reward Density} in deep search scenarios, where agents expend significant exploratory costs for infrequent and often null final rewards. In this paper, we formalize this challenge as the \textbf{Reward Density Optimization} problem, which aims to improve the reward obtained per unit of exploration cost. This paper introduce \textbf{InfoFlow}, a systematic framework that tackles this problem from three aspects. 1) \textbf{Subproblem decomposition}: breaking down long-range tasks to assign process rewards, thereby providing denser learning signals. 2) \textbf{Failure-guided hints}: injecting corrective guidance into stalled trajectories to increase the probability of successful outcomes. 3) \textbf{Dual-agent refinement}: employing a dual-agent architecture to offload the cognitive burden of deep exploration. A refiner agent synthesizes the search history, which effectively compresses the researcher's perceived trajectory, thereby reducing exploration cost and increasing the overall reward density. We evaluate InfoFlow on multiple agentic search benchmarks, where it significantly outperforms strong baselines, enabling lightweight LLMs to achieve performance comparable to advanced proprietary LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26575
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
Luo, Kun
Qian, Hongjin
Liu, Zheng
Xia, Ziyi
Xiao, Shitao
Bao, Siqi
Zhao, Jun
Liu, Kang
Computation and Language
Artificial Intelligence
Reinforcement Learning with Verifiable Rewards (RLVR) is a promising approach for enhancing agentic deep search. However, its application is often hindered by low \textbf{Reward Density} in deep search scenarios, where agents expend significant exploratory costs for infrequent and often null final rewards. In this paper, we formalize this challenge as the \textbf{Reward Density Optimization} problem, which aims to improve the reward obtained per unit of exploration cost. This paper introduce \textbf{InfoFlow}, a systematic framework that tackles this problem from three aspects. 1) \textbf{Subproblem decomposition}: breaking down long-range tasks to assign process rewards, thereby providing denser learning signals. 2) \textbf{Failure-guided hints}: injecting corrective guidance into stalled trajectories to increase the probability of successful outcomes. 3) \textbf{Dual-agent refinement}: employing a dual-agent architecture to offload the cognitive burden of deep exploration. A refiner agent synthesizes the search history, which effectively compresses the researcher's perceived trajectory, thereby reducing exploration cost and increasing the overall reward density. We evaluate InfoFlow on multiple agentic search benchmarks, where it significantly outperforms strong baselines, enabling lightweight LLMs to achieve performance comparable to advanced proprietary LLMs.
title InfoFlow: Reinforcing Search Agent Via Reward Density Optimization
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.26575