Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sorokin, Artyom, Buzun, Nazar, Anokhin, Alexander, Inozemcev, Oleg, Vedernikov, Egor, Anokhin, Petr, Burtsev, Mikhail, Alexey, Trushkov, Wenshuai, Yin, Burnaev, Evgeny
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911643683782656
author Sorokin, Artyom
Buzun, Nazar
Anokhin, Alexander
Inozemcev, Oleg
Vedernikov, Egor
Anokhin, Petr
Burtsev, Mikhail
Alexey, Trushkov
Wenshuai, Yin
Burnaev, Evgeny
author_facet Sorokin, Artyom
Buzun, Nazar
Anokhin, Alexander
Inozemcev, Oleg
Vedernikov, Egor
Anokhin, Petr
Burtsev, Mikhail
Alexey, Trushkov
Wenshuai, Yin
Burnaev, Evgeny
contents Retrieval-Augmented Generation (RAG) methods enhance LLM performance by efficiently filtering relevant context for LLMs, reducing hallucinations and inference cost. However, most existing RAG methods focus on single-step retrieval, which is often insufficient for answering complex questions that require multi-step search. Recently, multi-step retrieval approaches have emerged, typically involving the fine-tuning of small LLMs to perform multi-step retrieval. This type of fine-tuning is highly resource-intensive and does not enable the use of larger LLMs. In this work, we propose Q-RAG, a novel approach that fine-tunes the Embedder model for multi-step retrieval using reinforcement learning (RL). Q-RAG offers a competitive, resource-efficient alternative to existing multi-step retrieval methods for open-domain question answering and achieves state-of-the-art results on the popular long-context benchmarks BabiLong and RULER for contexts up to 10M tokens. Code is available at https://github.com/griver/Q-RAG
format Preprint
id arxiv_https___arxiv_org_abs_2511_07328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
Sorokin, Artyom
Buzun, Nazar
Anokhin, Alexander
Inozemcev, Oleg
Vedernikov, Egor
Anokhin, Petr
Burtsev, Mikhail
Alexey, Trushkov
Wenshuai, Yin
Burnaev, Evgeny
Machine Learning
Information Retrieval
Retrieval-Augmented Generation (RAG) methods enhance LLM performance by efficiently filtering relevant context for LLMs, reducing hallucinations and inference cost. However, most existing RAG methods focus on single-step retrieval, which is often insufficient for answering complex questions that require multi-step search. Recently, multi-step retrieval approaches have emerged, typically involving the fine-tuning of small LLMs to perform multi-step retrieval. This type of fine-tuning is highly resource-intensive and does not enable the use of larger LLMs. In this work, we propose Q-RAG, a novel approach that fine-tunes the Embedder model for multi-step retrieval using reinforcement learning (RL). Q-RAG offers a competitive, resource-efficient alternative to existing multi-step retrieval methods for open-domain question answering and achieves state-of-the-art results on the popular long-context benchmarks BabiLong and RULER for contexts up to 10M tokens. Code is available at https://github.com/griver/Q-RAG
title Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
topic Machine Learning
Information Retrieval
url https://arxiv.org/abs/2511.07328