Adapting Large Language Models for Multi-Domain Retrieval-Augmented-Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915225469452288 |
|---|---|
| author | Misrahi, Alexandre Chirkova, Nadezhda Louis, Maxime Nikoulina, Vassilina |
| author_facet | Misrahi, Alexandre Chirkova, Nadezhda Louis, Maxime Nikoulina, Vassilina |
| contents | Retrieval-Augmented Generation (RAG) enhances LLM factuality, but multi-domain applications face challenges like lack of diverse benchmarks and poor out-of-domain generalization. The first contribution of this work is to introduce a diverse benchmark comprising a variety of question-answering tasks from 8 sources and covering 13 domains. Our second contribution consists in systematically testing out-of-domain generalization for typical RAG tuning strategies. While our findings reveal that standard fine-tuning fails to generalize effectively, we show that sequence-level distillation with teacher-generated labels improves out-of-domain performance by providing more coherent supervision. Our findings highlight key strategies for improving multi-domain RAG robustness. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_02411 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Adapting Large Language Models for Multi-Domain Retrieval-Augmented-Generation Misrahi, Alexandre Chirkova, Nadezhda Louis, Maxime Nikoulina, Vassilina Computation and Language Retrieval-Augmented Generation (RAG) enhances LLM factuality, but multi-domain applications face challenges like lack of diverse benchmarks and poor out-of-domain generalization. The first contribution of this work is to introduce a diverse benchmark comprising a variety of question-answering tasks from 8 sources and covering 13 domains. Our second contribution consists in systematically testing out-of-domain generalization for typical RAG tuning strategies. While our findings reveal that standard fine-tuning fails to generalize effectively, we show that sequence-level distillation with teacher-generated labels improves out-of-domain performance by providing more coherent supervision. Our findings highlight key strategies for improving multi-domain RAG robustness. |
| title | Adapting Large Language Models for Multi-Domain Retrieval-Augmented-Generation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2504.02411 |