An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Bowen, Yoon, Jinsung, Kargupta, Priyanka, Arik, Sercan O., Han, Jiawei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915295747112960
author Jin, Bowen
Yoon, Jinsung
Kargupta, Priyanka
Arik, Sercan O.
Han, Jiawei
author_facet Jin, Bowen
Yoon, Jinsung
Kargupta, Priyanka
Arik, Sercan O.
Han, Jiawei
contents Reinforcement learning (RL) has demonstrated strong potential in training large language models (LLMs) capable of complex reasoning for real-world problem solving. More recently, RL has been leveraged to create sophisticated LLM-based search agents that adeptly combine reasoning with search engine use. While the use of RL for training search agents is promising, the optimal design of such agents remains not fully understood. In particular, key factors -- such as (1) reward formulation, (2) the choice and characteristics of the underlying LLM, and (3) the role of the search engine in the RL process -- require further investigation. In this work, we conduct comprehensive empirical studies to systematically investigate these and offer actionable insights. We highlight several key findings: format rewards are effective in improving final performance, whereas intermediate retrieval rewards have limited impact; the scale and initialization of the LLM (general-purpose vs. reasoning-specialized) significantly influence RL outcomes; and the choice of search engine plays a critical role in shaping RL training dynamics and the robustness of the trained agent during inference. These establish important guidelines for successfully building and deploying LLM-based search agents in real-world applications. Code is available at https://github.com/PeterGriffinJin/Search-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15117
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
Jin, Bowen
Yoon, Jinsung
Kargupta, Priyanka
Arik, Sercan O.
Han, Jiawei
Computation and Language
Artificial Intelligence
Information Retrieval
Reinforcement learning (RL) has demonstrated strong potential in training large language models (LLMs) capable of complex reasoning for real-world problem solving. More recently, RL has been leveraged to create sophisticated LLM-based search agents that adeptly combine reasoning with search engine use. While the use of RL for training search agents is promising, the optimal design of such agents remains not fully understood. In particular, key factors -- such as (1) reward formulation, (2) the choice and characteristics of the underlying LLM, and (3) the role of the search engine in the RL process -- require further investigation. In this work, we conduct comprehensive empirical studies to systematically investigate these and offer actionable insights. We highlight several key findings: format rewards are effective in improving final performance, whereas intermediate retrieval rewards have limited impact; the scale and initialization of the LLM (general-purpose vs. reasoning-specialized) significantly influence RL outcomes; and the choice of search engine plays a critical role in shaping RL training dynamics and the robustness of the trained agent during inference. These establish important guidelines for successfully building and deploying LLM-based search agents in real-world applications. Code is available at https://github.com/PeterGriffinJin/Search-R1.
title An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2505.15117