LSTM-based Selective Dense Text Retrieval Guided by Sparse Lexical Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Yingrui, Carlson, Parker, Qiao, Yifan, Xie, Wentai, He, Shanxiu, Yang, Tao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912233284435968
author Yang, Yingrui
Carlson, Parker
Qiao, Yifan
Xie, Wentai
He, Shanxiu
Yang, Tao
author_facet Yang, Yingrui
Carlson, Parker
Qiao, Yifan
Xie, Wentai
He, Shanxiu
Yang, Tao
contents This paper studies fast fusion of dense retrieval and sparse lexical retrieval, and proposes a cluster-based selective dense retrieval method called CluSD guided by sparse lexical retrieval. CluSD takes a lightweight cluster-based approach and exploits the overlap of sparse retrieval results and embedding clusters in a two-stage selection process with an LSTM model to quickly identify relevant clusters while incurring limited extra memory space overhead. CluSD triggers partial dense retrieval and performs cluster-based block disk I/O if needed. This paper evaluates CluSD and compares it with several baselines for searching in-memory and on-disk MS MARCO and BEIR datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2502_10639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LSTM-based Selective Dense Text Retrieval Guided by Sparse Lexical Retrieval
Yang, Yingrui
Carlson, Parker
Qiao, Yifan
Xie, Wentai
He, Shanxiu
Yang, Tao
Information Retrieval
This paper studies fast fusion of dense retrieval and sparse lexical retrieval, and proposes a cluster-based selective dense retrieval method called CluSD guided by sparse lexical retrieval. CluSD takes a lightweight cluster-based approach and exploits the overlap of sparse retrieval results and embedding clusters in a two-stage selection process with an LSTM model to quickly identify relevant clusters while incurring limited extra memory space overhead. CluSD triggers partial dense retrieval and performs cluster-based block disk I/O if needed. This paper evaluates CluSD and compares it with several baselines for searching in-memory and on-disk MS MARCO and BEIR datasets.
title LSTM-based Selective Dense Text Retrieval Guided by Sparse Lexical Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2502.10639