WARP: An Efficient Engine for Multi-Vector Retrieval

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Scheerer, Jan Luca, Zaharia, Matei, Potts, Christopher, Alonso, Gustavo, Khattab, Omar
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916828194799616
author Scheerer, Jan Luca
Zaharia, Matei
Potts, Christopher
Alonso, Gustavo
Khattab, Omar
author_facet Scheerer, Jan Luca
Zaharia, Matei
Potts, Christopher
Alonso, Gustavo
Khattab, Omar
contents Multi-vector retrieval methods such as ColBERT and its recent variant, the ConteXtualized Token Retriever (XTR), offer high accuracy but face efficiency challenges at scale. To address this, we present WARP, a retrieval engine that substantially improves the efficiency of retrievers trained with the XTR objective through three key innovations: (1) WARP$_\text{SELECT}$ for dynamic similarity imputation; (2) implicit decompression, avoiding costly vector reconstruction during retrieval; and (3) a two-stage reduction process for efficient score aggregation. Combined with highly-optimized C++ kernels, our system reduces end-to-end latency compared to XTR's reference implementation by 41x, and achieves a 3x speedup over the ColBERTv2/PLAID engine, while preserving retrieval quality.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17788
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WARP: An Efficient Engine for Multi-Vector Retrieval
Scheerer, Jan Luca
Zaharia, Matei
Potts, Christopher
Alonso, Gustavo
Khattab, Omar
Information Retrieval
Multi-vector retrieval methods such as ColBERT and its recent variant, the ConteXtualized Token Retriever (XTR), offer high accuracy but face efficiency challenges at scale. To address this, we present WARP, a retrieval engine that substantially improves the efficiency of retrievers trained with the XTR objective through three key innovations: (1) WARP$_\text{SELECT}$ for dynamic similarity imputation; (2) implicit decompression, avoiding costly vector reconstruction during retrieval; and (3) a two-stage reduction process for efficient score aggregation. Combined with highly-optimized C++ kernels, our system reduces end-to-end latency compared to XTR's reference implementation by 41x, and achieves a 3x speedup over the ColBERTv2/PLAID engine, while preserving retrieval quality.
title WARP: An Efficient Engine for Multi-Vector Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2501.17788