MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kalra, Jushaan Singh, Zhao, Xinran, Kim, To Eun, Cai, Fengyu, Diaz, Fernando, Wu, Tongshuang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916800246054912
author Kalra, Jushaan Singh
Zhao, Xinran
Kim, To Eun
Cai, Fengyu
Diaz, Fernando
Wu, Tongshuang
author_facet Kalra, Jushaan Singh
Zhao, Xinran
Kim, To Eun
Cai, Fengyu
Diaz, Fernando
Wu, Tongshuang
contents Retrieval-augmented Generation (RAG) is powerful, but its effectiveness hinges on which retrievers we use and how. Different retrievers offer distinct, often complementary signals: BM25 captures lexical matches; dense retrievers, semantic similarity. Yet in practice, we typically fix a single retriever based on heuristics, which fails to generalize across diverse information needs. Can we dynamically select and integrate multiple retrievers for each individual query, without the need for manual selection? In our work, we validate this intuition with quantitative analysis and introduce mixture of retrievers: a zero-shot, weighted combination of heterogeneous retrievers. Extensive experiments show that such mixtures are effective and efficient: Despite totaling just 0.8B parameters, this mixture outperforms every individual retriever and even larger 7B models by +10.8% and +3.9% on average, respectively. Further analysis also shows that this mixture framework can help incorporate specialized non-oracle human information sources as retrievers to achieve good collaboration, with a 58.9% relative performance improvement over simulated humans alone.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15862
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
Kalra, Jushaan Singh
Zhao, Xinran
Kim, To Eun
Cai, Fengyu
Diaz, Fernando
Wu, Tongshuang
Information Retrieval
Artificial Intelligence
Computation and Language
Retrieval-augmented Generation (RAG) is powerful, but its effectiveness hinges on which retrievers we use and how. Different retrievers offer distinct, often complementary signals: BM25 captures lexical matches; dense retrievers, semantic similarity. Yet in practice, we typically fix a single retriever based on heuristics, which fails to generalize across diverse information needs. Can we dynamically select and integrate multiple retrievers for each individual query, without the need for manual selection? In our work, we validate this intuition with quantitative analysis and introduce mixture of retrievers: a zero-shot, weighted combination of heterogeneous retrievers. Extensive experiments show that such mixtures are effective and efficient: Despite totaling just 0.8B parameters, this mixture outperforms every individual retriever and even larger 7B models by +10.8% and +3.9% on average, respectively. Further analysis also shows that this mixture framework can help incorporate specialized non-oracle human information sources as retrievers to achieve good collaboration, with a 58.9% relative performance improvement over simulated humans alone.
title MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.15862