AGRaME: Any-Granularity Ranking with Multi-Vector Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Reddy, Revanth Gangi, Attia, Omar, Li, Yunyao, Ji, Heng, Potdar, Saloni
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916258842148864
author Reddy, Revanth Gangi
Attia, Omar
Li, Yunyao
Ji, Heng
Potdar, Saloni
author_facet Reddy, Revanth Gangi
Attia, Omar
Li, Yunyao
Ji, Heng
Potdar, Saloni
contents Ranking is a fundamental and popular problem in search. However, existing ranking algorithms usually restrict the granularity of ranking to full passages or require a specific dense index for each desired level of granularity. Such lack of flexibility in granularity negatively affects many applications that can benefit from more granular ranking, such as sentence-level ranking for open-domain question-answering, or proposition-level ranking for attribution. In this work, we introduce the idea of any-granularity ranking, which leverages multi-vector embeddings to rank at varying levels of granularity while maintaining encoding at a single (coarser) level of granularity. We propose a multi-granular contrastive loss for training multi-vector approaches, and validate its utility with both sentences and propositions as ranking units. Finally, we demonstrate the application of proposition-level ranking to post-hoc citation addition in retrieval-augmented generation, surpassing the performance of prompt-driven citation generation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15028
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AGRaME: Any-Granularity Ranking with Multi-Vector Embeddings
Reddy, Revanth Gangi
Attia, Omar
Li, Yunyao
Ji, Heng
Potdar, Saloni
Computation and Language
Information Retrieval
Ranking is a fundamental and popular problem in search. However, existing ranking algorithms usually restrict the granularity of ranking to full passages or require a specific dense index for each desired level of granularity. Such lack of flexibility in granularity negatively affects many applications that can benefit from more granular ranking, such as sentence-level ranking for open-domain question-answering, or proposition-level ranking for attribution. In this work, we introduce the idea of any-granularity ranking, which leverages multi-vector embeddings to rank at varying levels of granularity while maintaining encoding at a single (coarser) level of granularity. We propose a multi-granular contrastive loss for training multi-vector approaches, and validate its utility with both sentences and propositions as ranking units. Finally, we demonstrate the application of proposition-level ranking to post-hoc citation addition in retrieval-augmented generation, surpassing the performance of prompt-driven citation generation.
title AGRaME: Any-Granularity Ranking with Multi-Vector Embeddings
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2405.15028