Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Sheryl, Khattab, Omar, Finn, Chelsea, Sharma, Archit
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913568343982080
author Hsu, Sheryl
Khattab, Omar
Finn, Chelsea
Sharma, Archit
author_facet Hsu, Sheryl
Khattab, Omar
Finn, Chelsea
Sharma, Archit
contents The hallucinations of large language models (LLMs) are increasingly mitigated by allowing LLMs to search for information and to ground their answers in real sources. Unfortunately, LLMs often struggle with posing the right search queries, especially when dealing with complex or otherwise indirect topics. Observing that LLMs can learn to search for relevant facts by $\textit{trying}$ different queries and learning to up-weight queries that successfully produce relevant results, we introduce $\underline{Le}$arning to $\underline{Re}$trieve by $\underline{T}$rying (LeReT), a reinforcement learning framework that explores search queries and uses preference-based optimization to improve their quality. LeReT can improve the absolute retrieval accuracy by up to 29% and the downstream generator evaluations by 17%. The simplicity and flexibility of LeReT allows it to be applied to arbitrary off-the-shelf retrievers and makes it a promising technique for improving general LLM pipelines. Project website: http://sherylhsu.com/LeReT/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23214
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
Hsu, Sheryl
Khattab, Omar
Finn, Chelsea
Sharma, Archit
Machine Learning
Artificial Intelligence
The hallucinations of large language models (LLMs) are increasingly mitigated by allowing LLMs to search for information and to ground their answers in real sources. Unfortunately, LLMs often struggle with posing the right search queries, especially when dealing with complex or otherwise indirect topics. Observing that LLMs can learn to search for relevant facts by $\textit{trying}$ different queries and learning to up-weight queries that successfully produce relevant results, we introduce $\underline{Le}$arning to $\underline{Re}$trieve by $\underline{T}$rying (LeReT), a reinforcement learning framework that explores search queries and uses preference-based optimization to improve their quality. LeReT can improve the absolute retrieval accuracy by up to 29% and the downstream generator evaluations by 17%. The simplicity and flexibility of LeReT allows it to be applied to arbitrary off-the-shelf retrievers and makes it a promising technique for improving general LLM pipelines. Project website: http://sherylhsu.com/LeReT/.
title Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.23214