Approximate Cluster-Based Sparse Document Retrieval with Segmented Maximum Term Weights

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qiao, Yifan, He, Shanxiu, Yang, Yingrui, Carlson, Parker, Yang, Tao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929313251590144
author Qiao, Yifan
He, Shanxiu
Yang, Yingrui
Carlson, Parker
Yang, Tao
author_facet Qiao, Yifan
He, Shanxiu
Yang, Yingrui
Carlson, Parker
Yang, Tao
contents This paper revisits cluster-based retrieval that partitions the inverted index into multiple groups and skips the index partially at cluster and document levels during online inference using a learned sparse representation. It proposes an approximate search scheme with two parameters to control the rank-safeness competitiveness of pruning with segmented maximum term weights within each cluster. Cluster-level maximum weight segmentation allows an improvement in the rank score bound estimation and threshold-based pruning to be approximately adaptive to bound estimation tightness, resulting in better relevance and efficiency. The experiments with MS MARCO passage ranking and BEIR datasets demonstrate the usefulness of the proposed scheme with a comparison to the baselines. This paper presents the design of this approximate retrieval scheme with rank-safeness analysis, compares clustering and segmentation options, and reports evaluation results.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08896
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Approximate Cluster-Based Sparse Document Retrieval with Segmented Maximum Term Weights
Qiao, Yifan
He, Shanxiu
Yang, Yingrui
Carlson, Parker
Yang, Tao
Information Retrieval
This paper revisits cluster-based retrieval that partitions the inverted index into multiple groups and skips the index partially at cluster and document levels during online inference using a learned sparse representation. It proposes an approximate search scheme with two parameters to control the rank-safeness competitiveness of pruning with segmented maximum term weights within each cluster. Cluster-level maximum weight segmentation allows an improvement in the rank score bound estimation and threshold-based pruning to be approximately adaptive to bound estimation tightness, resulting in better relevance and efficiency. The experiments with MS MARCO passage ranking and BEIR datasets demonstrate the usefulness of the proposed scheme with a comparison to the baselines. This paper presents the design of this approximate retrieval scheme with rank-safeness analysis, compares clustering and segmentation options, and reports evaluation results.
title Approximate Cluster-Based Sparse Document Retrieval with Segmented Maximum Term Weights
topic Information Retrieval
url https://arxiv.org/abs/2404.08896