Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Wenlin, Li, Xiangyang, Dong, Kuicai, Wang, Yichao, Jia, Pengyue, Li, Xiaopeng, Zhang, Yingyi, Xu, Derong, Du, Zhaocheng, Guo, Huifeng, Tang, Ruiming, Zhao, Xiangyu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909946963034112
author Zhang, Wenlin
Li, Xiangyang
Dong, Kuicai
Wang, Yichao
Jia, Pengyue
Li, Xiaopeng
Zhang, Yingyi
Xu, Derong
Du, Zhaocheng
Guo, Huifeng
Tang, Ruiming
Zhao, Xiangyu
author_facet Zhang, Wenlin
Li, Xiangyang
Dong, Kuicai
Wang, Yichao
Jia, Pengyue
Li, Xiaopeng
Zhang, Yingyi
Xu, Derong
Du, Zhaocheng
Guo, Huifeng
Tang, Ruiming
Zhao, Xiangyu
contents Retrieval-augmented generation (RAG) enhances the text generation capabilities of large language models (LLMs) by integrating external knowledge and up-to-date information. However, traditional RAG systems are limited by static workflows and lack the adaptability required for multistep reasoning and complex task management. To address these limitations, agentic RAG systems (e.g., DeepResearch) have been proposed, enabling dynamic retrieval strategies, iterative context refinement, and adaptive workflows for handling complex search queries beyond the capabilities of conventional RAG. Recent advances, such as Search-R1, have demonstrated promising gains using outcome-based reinforcement learning, where the correctness of the final answer serves as the reward signal. Nevertheless, such outcome-supervised agentic RAG methods face challenges including low exploration efficiency, gradient conflict, and sparse reward signals. To overcome these challenges, we propose to utilize fine-grained, process-level rewards to improve training stability, reduce computational costs, and enhance efficiency. Specifically, we introduce a novel method ReasonRAG that automatically constructs RAG-ProGuide, a high-quality dataset providing process-level rewards for (i) query generation, (ii) evidence extraction, and (iii) answer generation, thereby enhancing model inherent capabilities via process-supervised reinforcement learning. With the process-level policy optimization, the proposed framework empowers LLMs to autonomously invoke search, generate queries, extract relevant evidence, and produce final answers. Compared to existing approaches such as Search-R1 and traditional RAG systems, ReasonRAG, leveraging RAG-ProGuide, achieves superior performance on five benchmark datasets using only 5k training instances, significantly fewer than the 90k training instances required by Search-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14069
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning
Zhang, Wenlin
Li, Xiangyang
Dong, Kuicai
Wang, Yichao
Jia, Pengyue
Li, Xiaopeng
Zhang, Yingyi
Xu, Derong
Du, Zhaocheng
Guo, Huifeng
Tang, Ruiming
Zhao, Xiangyu
Information Retrieval
Retrieval-augmented generation (RAG) enhances the text generation capabilities of large language models (LLMs) by integrating external knowledge and up-to-date information. However, traditional RAG systems are limited by static workflows and lack the adaptability required for multistep reasoning and complex task management. To address these limitations, agentic RAG systems (e.g., DeepResearch) have been proposed, enabling dynamic retrieval strategies, iterative context refinement, and adaptive workflows for handling complex search queries beyond the capabilities of conventional RAG. Recent advances, such as Search-R1, have demonstrated promising gains using outcome-based reinforcement learning, where the correctness of the final answer serves as the reward signal. Nevertheless, such outcome-supervised agentic RAG methods face challenges including low exploration efficiency, gradient conflict, and sparse reward signals. To overcome these challenges, we propose to utilize fine-grained, process-level rewards to improve training stability, reduce computational costs, and enhance efficiency. Specifically, we introduce a novel method ReasonRAG that automatically constructs RAG-ProGuide, a high-quality dataset providing process-level rewards for (i) query generation, (ii) evidence extraction, and (iii) answer generation, thereby enhancing model inherent capabilities via process-supervised reinforcement learning. With the process-level policy optimization, the proposed framework empowers LLMs to autonomously invoke search, generate queries, extract relevant evidence, and produce final answers. Compared to existing approaches such as Search-R1 and traditional RAG systems, ReasonRAG, leveraging RAG-ProGuide, achieves superior performance on five benchmark datasets using only 5k training instances, significantly fewer than the 90k training instances required by Search-R1.
title Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning
topic Information Retrieval
url https://arxiv.org/abs/2505.14069