SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jingsheng, Li, Linxu, Li, Weiyuan, Fu, Yuzhuo, Dai, Bin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913726478680064
author Gao, Jingsheng
Li, Linxu
Li, Weiyuan
Fu, Yuzhuo
Dai, Bin
author_facet Gao, Jingsheng
Li, Linxu
Li, Weiyuan
Fu, Yuzhuo
Dai, Bin
contents RAG systems consist of multiple modules to work together. However, these modules are usually separately trained. We argue that a system like RAG that incorporates multiple modules should be jointly optimized to achieve optimal performance. To demonstrate this, we design a specific pipeline called \textbf{SmartRAG} that includes a policy network and a retriever. The policy network can serve as 1) a decision maker that decides when to retrieve, 2) a query rewriter to generate a query most suited to the retriever, and 3) an answer generator that produces the final response with/without the observations. We then propose to jointly optimize the whole system using a reinforcement learning algorithm, with the reward designed to encourage the system to achieve the best performance with minimal retrieval cost. When jointly optimized, all the modules can be aware of how other modules are working and thus find the best way to work together as a complete system. Empirical results demonstrate that the jointly optimized SmartRAG can achieve better performance than separately optimized counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18141
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback
Gao, Jingsheng
Li, Linxu
Li, Weiyuan
Fu, Yuzhuo
Dai, Bin
Information Retrieval
Artificial Intelligence
Computation and Language
RAG systems consist of multiple modules to work together. However, these modules are usually separately trained. We argue that a system like RAG that incorporates multiple modules should be jointly optimized to achieve optimal performance. To demonstrate this, we design a specific pipeline called \textbf{SmartRAG} that includes a policy network and a retriever. The policy network can serve as 1) a decision maker that decides when to retrieve, 2) a query rewriter to generate a query most suited to the retriever, and 3) an answer generator that produces the final response with/without the observations. We then propose to jointly optimize the whole system using a reinforcement learning algorithm, with the reward designed to encourage the system to achieve the best performance with minimal retrieval cost. When jointly optimized, all the modules can be aware of how other modules are working and thus find the best way to work together as a complete system. Empirical results demonstrate that the jointly optimized SmartRAG can achieve better performance than separately optimized counterparts.
title SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.18141