s3: You Don't Need That Much Data to Train a Search Agent via RL

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jiang, Pengcheng, Xu, Xueqiang, Lin, Jiacheng, Xiao, Jinfeng, Wang, Zifeng, Sun, Jimeng, Han, Jiawei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914136287346688
author Jiang, Pengcheng
Xu, Xueqiang
Lin, Jiacheng
Xiao, Jinfeng
Wang, Zifeng
Sun, Jimeng
Han, Jiawei
author_facet Jiang, Pengcheng
Xu, Xueqiang
Lin, Jiacheng
Xiao, Jinfeng
Wang, Zifeng
Sun, Jimeng
Han, Jiawei
contents Retrieval-augmented generation (RAG) systems empower large language models (LLMs) to access external knowledge during inference. Recent advances have enabled LLMs to act as search agents via reinforcement learning (RL), improving information acquisition through multi-turn interactions with retrieval engines. However, existing approaches either optimize retrieval using search-only metrics (e.g., NDCG) that ignore downstream utility or fine-tune the entire LLM to jointly reason and retrieve-entangling retrieval with generation and limiting the real search utility and compatibility with frozen or proprietary models. In this work, we propose s3, a lightweight, model-agnostic framework that decouples the searcher from the generator and trains the searcher using a Gain Beyond RAG reward: the improvement in generation accuracy over naive RAG. s3 requires only 2.4k training samples to outperform baselines trained on over 70x more data, consistently delivering stronger downstream performance across six general QA and five medical QA benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14146
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle s3: You Don't Need That Much Data to Train a Search Agent via RL
Jiang, Pengcheng
Xu, Xueqiang
Lin, Jiacheng
Xiao, Jinfeng
Wang, Zifeng
Sun, Jimeng
Han, Jiawei
Artificial Intelligence
Computation and Language
Retrieval-augmented generation (RAG) systems empower large language models (LLMs) to access external knowledge during inference. Recent advances have enabled LLMs to act as search agents via reinforcement learning (RL), improving information acquisition through multi-turn interactions with retrieval engines. However, existing approaches either optimize retrieval using search-only metrics (e.g., NDCG) that ignore downstream utility or fine-tune the entire LLM to jointly reason and retrieve-entangling retrieval with generation and limiting the real search utility and compatibility with frozen or proprietary models. In this work, we propose s3, a lightweight, model-agnostic framework that decouples the searcher from the generator and trains the searcher using a Gain Beyond RAG reward: the improvement in generation accuracy over naive RAG. s3 requires only 2.4k training samples to outperform baselines trained on over 70x more data, consistently delivering stronger downstream performance across six general QA and five medical QA benchmarks.
title s3: You Don't Need That Much Data to Train a Search Agent via RL
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.14146