Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Hao, Xie, Peitong, Chen, Jingxue, Lin, Jie, Tang, Qingkun, Lu, Qianchun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916939801034752
author Lin, Hao
Xie, Peitong
Chen, Jingxue
Lin, Jie
Tang, Qingkun
Lu, Qianchun
author_facet Lin, Hao
Xie, Peitong
Chen, Jingxue
Lin, Jie
Tang, Qingkun
Lu, Qianchun
contents Retrieval-Augmented Generation (RAG) systems rely heavily on the retrieval stage, particularly the coarse-ranking process. Existing coarse-ranking optimization approaches often struggle to balance domain-specific knowledge learning with query enhencement, resulting in suboptimal retrieval performance. To address this challenge, we propose MoLER, a domain-aware RAG method that uses MoL-Enhanced Reinforcement Learning to optimize retrieval. MoLER has a two-stage pipeline: a continual pre-training (CPT) phase using a Mixture of Losses (MoL) to balance domain-specific knowledge with general language capabilities, and a reinforcement learning (RL) phase leveraging Group Relative Policy Optimization (GRPO) to optimize query and passage generation for maximizing document recall. A key innovation is our Multi-query Single-passage Late Fusion (MSLF) strategy, which reduces computational overhead during RL training while maintaining scalable inference via Multi-query Multi-passage Late Fusion (MMLF). Extensive experiments on benchmark datasets show that MoLER achieves state-of-the-art performance, significantly outperforming baseline methods. MoLER bridges the knowledge gap in RAG systems, enabling robust and scalable retrieval in specialized domains.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval
Lin, Hao
Xie, Peitong
Chen, Jingxue
Lin, Jie
Tang, Qingkun
Lu, Qianchun
Computation and Language
Information Retrieval
Retrieval-Augmented Generation (RAG) systems rely heavily on the retrieval stage, particularly the coarse-ranking process. Existing coarse-ranking optimization approaches often struggle to balance domain-specific knowledge learning with query enhencement, resulting in suboptimal retrieval performance. To address this challenge, we propose MoLER, a domain-aware RAG method that uses MoL-Enhanced Reinforcement Learning to optimize retrieval. MoLER has a two-stage pipeline: a continual pre-training (CPT) phase using a Mixture of Losses (MoL) to balance domain-specific knowledge with general language capabilities, and a reinforcement learning (RL) phase leveraging Group Relative Policy Optimization (GRPO) to optimize query and passage generation for maximizing document recall. A key innovation is our Multi-query Single-passage Late Fusion (MSLF) strategy, which reduces computational overhead during RL training while maintaining scalable inference via Multi-query Multi-passage Late Fusion (MMLF). Extensive experiments on benchmark datasets show that MoLER achieves state-of-the-art performance, significantly outperforming baseline methods. MoLER bridges the knowledge gap in RAG systems, enabling robust and scalable retrieval in specialized domains.
title Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2509.06650