Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cao, Hongyu, Wu, Yuxuan, Cai, Yucheng, Zhao, Xianyu, Ou, Zhijian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908986343686144
author Cao, Hongyu
Wu, Yuxuan
Cai, Yucheng
Zhao, Xianyu
Ou, Zhijian
author_facet Cao, Hongyu
Wu, Yuxuan
Cai, Yucheng
Zhao, Xianyu
Ou, Zhijian
contents Retrieval-augmented generation (RAG) has become a widely recognized paradigm to combine parametric memory with non-parametric memories. An RAG model consists of two serial connecting components (retriever and generator). A major challenge in end-to-end optimization of the RAG model is that marginalization over relevant passages (modeled as discrete latent variables) from a knowledge base is required. Traditional top-K marginalization and variational RAG (VRAG) suffer from biased or high-variance gradient estimates. In this paper, we propose and develop joint stochastic approximation (JSA) based end-to-end training of RAG, which is referred to as JSA-RAG. The JSA algorithm is a stochastic extension of the EM (expectation-maximization) algorithm and is particularly powerful in estimating discrete latent variable models. Extensive experiments are conducted on five datasets for two tasks (open-domain question answering, knowledge-grounded dialogs) and show that JSA-RAG significantly outperforms both vanilla RAG and VRAG. Further analysis shows the efficacy of JSA-RAG from the perspectives of generation, retrieval, and low-variance gradient estimate.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18168
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation
Cao, Hongyu
Wu, Yuxuan
Cai, Yucheng
Zhao, Xianyu
Ou, Zhijian
Computation and Language
Retrieval-augmented generation (RAG) has become a widely recognized paradigm to combine parametric memory with non-parametric memories. An RAG model consists of two serial connecting components (retriever and generator). A major challenge in end-to-end optimization of the RAG model is that marginalization over relevant passages (modeled as discrete latent variables) from a knowledge base is required. Traditional top-K marginalization and variational RAG (VRAG) suffer from biased or high-variance gradient estimates. In this paper, we propose and develop joint stochastic approximation (JSA) based end-to-end training of RAG, which is referred to as JSA-RAG. The JSA algorithm is a stochastic extension of the EM (expectation-maximization) algorithm and is particularly powerful in estimating discrete latent variable models. Extensive experiments are conducted on five datasets for two tasks (open-domain question answering, knowledge-grounded dialogs) and show that JSA-RAG significantly outperforms both vanilla RAG and VRAG. Further analysis shows the efficacy of JSA-RAG from the perspectives of generation, retrieval, and low-variance gradient estimate.
title Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation
topic Computation and Language
url https://arxiv.org/abs/2508.18168