Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hung-Ting, Liu, Xiang, Ravfogel, Shauli, Choi, Eunsol
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912687900852224
author Chen, Hung-Ting
Liu, Xiang
Ravfogel, Shauli
Choi, Eunsol
author_facet Chen, Hung-Ting
Liu, Xiang
Ravfogel, Shauli
Choi, Eunsol
contents Most text retrievers generate \emph{one} query vector to retrieve relevant documents. Yet, the conditional distribution of relevant documents for the query may be multimodal, e.g., representing different interpretations of the query. We first quantify the limitations of existing retrievers. All retrievers we evaluate struggle more as the distance between target document embeddings grows. To address this limitation, we develop a new retriever architecture, \emph{A}utoregressive \emph{M}ulti-\emph{E}mbedding \emph{R}etriever (AMER). Our model autoregressively generates multiple query vectors, and all the predicted query vectors are used to retrieve documents from the corpus. We show that on the synthetic vectorized data, the proposed method could capture multiple target distributions perfectly, showing 4x better performance than single embedding model. We also fine-tune our model on real-world multi-answer retrieval datasets and evaluate in-domain. AMER presents 4 and 21\% relative gains over single-embedding baselines on two datasets we evaluate on. Furthermore, we consistently observe larger gains on the subset of dataset where the embeddings of the target documents are less similar to each other. We demonstrate the potential of using a multi-query vector retriever and open up a new direction for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2511_02770
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
Chen, Hung-Ting
Liu, Xiang
Ravfogel, Shauli
Choi, Eunsol
Computation and Language
Information Retrieval
Most text retrievers generate \emph{one} query vector to retrieve relevant documents. Yet, the conditional distribution of relevant documents for the query may be multimodal, e.g., representing different interpretations of the query. We first quantify the limitations of existing retrievers. All retrievers we evaluate struggle more as the distance between target document embeddings grows. To address this limitation, we develop a new retriever architecture, \emph{A}utoregressive \emph{M}ulti-\emph{E}mbedding \emph{R}etriever (AMER). Our model autoregressively generates multiple query vectors, and all the predicted query vectors are used to retrieve documents from the corpus. We show that on the synthetic vectorized data, the proposed method could capture multiple target distributions perfectly, showing 4x better performance than single embedding model. We also fine-tune our model on real-world multi-answer retrieval datasets and evaluate in-domain. AMER presents 4 and 21\% relative gains over single-embedding baselines on two datasets we evaluate on. Furthermore, we consistently observe larger gains on the subset of dataset where the embeddings of the target documents are less similar to each other. We demonstrate the potential of using a multi-query vector retriever and open up a new direction for future work.
title Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2511.02770