SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Ran, Liu, Hui, Nag, Sreyashi, Dai, Zhenwei, Xie, Yaochen, Tang, Xianfeng, Luo, Chen, Li, Yang, Ho, Joyce C., Yang, Carl, He, Qi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910798912159744
author Xu, Ran
Liu, Hui
Nag, Sreyashi
Dai, Zhenwei
Xie, Yaochen
Tang, Xianfeng
Luo, Chen
Li, Yang
Ho, Joyce C.
Yang, Carl
He, Qi
author_facet Xu, Ran
Liu, Hui
Nag, Sreyashi
Dai, Zhenwei
Xie, Yaochen
Tang, Xianfeng
Luo, Chen
Li, Yang
Ho, Joyce C.
Yang, Carl
He, Qi
contents Retrieval-augmented generation (RAG) enhances the question-answering (QA) abilities of large language models (LLMs) by integrating external knowledge. However, adapting general-purpose RAG systems to specialized fields such as science and medicine poses unique challenges due to distribution shifts and limited access to domain-specific data. To tackle this, we propose SimRAG, a self-training approach that equips the LLM with joint capabilities of question answering and question generation for domain adaptation. Our method first fine-tunes the LLM on instruction-following, question-answering, and search-related data. Then, it prompts the same LLM to generate diverse domain-relevant questions from unlabeled corpora, with an additional filtering strategy to retain high-quality synthetic examples. By leveraging these self-generated synthetic examples, the LLM can improve their performance on domain-specific RAG tasks. Experiments on 11 datasets, spanning two backbone sizes and three domains, demonstrate that SimRAG outperforms baselines by 1.2\%--8.6\%.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17952
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains
Xu, Ran
Liu, Hui
Nag, Sreyashi
Dai, Zhenwei
Xie, Yaochen
Tang, Xianfeng
Luo, Chen
Li, Yang
Ho, Joyce C.
Yang, Carl
He, Qi
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
Retrieval-augmented generation (RAG) enhances the question-answering (QA) abilities of large language models (LLMs) by integrating external knowledge. However, adapting general-purpose RAG systems to specialized fields such as science and medicine poses unique challenges due to distribution shifts and limited access to domain-specific data. To tackle this, we propose SimRAG, a self-training approach that equips the LLM with joint capabilities of question answering and question generation for domain adaptation. Our method first fine-tunes the LLM on instruction-following, question-answering, and search-related data. Then, it prompts the same LLM to generate diverse domain-relevant questions from unlabeled corpora, with an additional filtering strategy to retain high-quality synthetic examples. By leveraging these self-generated synthetic examples, the LLM can improve their performance on domain-specific RAG tasks. Experiments on 11 datasets, spanning two backbone sizes and three domains, demonstrate that SimRAG outperforms baselines by 1.2\%--8.6\%.
title SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2410.17952