AGRaME: Any-Granularity Ranking with Multi-Vector Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916258842148864 |
|---|---|
| author | Reddy, Revanth Gangi Attia, Omar Li, Yunyao Ji, Heng Potdar, Saloni |
| author_facet | Reddy, Revanth Gangi Attia, Omar Li, Yunyao Ji, Heng Potdar, Saloni |
| contents | Ranking is a fundamental and popular problem in search. However, existing ranking algorithms usually restrict the granularity of ranking to full passages or require a specific dense index for each desired level of granularity. Such lack of flexibility in granularity negatively affects many applications that can benefit from more granular ranking, such as sentence-level ranking for open-domain question-answering, or proposition-level ranking for attribution. In this work, we introduce the idea of any-granularity ranking, which leverages multi-vector embeddings to rank at varying levels of granularity while maintaining encoding at a single (coarser) level of granularity. We propose a multi-granular contrastive loss for training multi-vector approaches, and validate its utility with both sentences and propositions as ranking units. Finally, we demonstrate the application of proposition-level ranking to post-hoc citation addition in retrieval-augmented generation, surpassing the performance of prompt-driven citation generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_15028 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | AGRaME: Any-Granularity Ranking with Multi-Vector Embeddings Reddy, Revanth Gangi Attia, Omar Li, Yunyao Ji, Heng Potdar, Saloni Computation and Language Information Retrieval Ranking is a fundamental and popular problem in search. However, existing ranking algorithms usually restrict the granularity of ranking to full passages or require a specific dense index for each desired level of granularity. Such lack of flexibility in granularity negatively affects many applications that can benefit from more granular ranking, such as sentence-level ranking for open-domain question-answering, or proposition-level ranking for attribution. In this work, we introduce the idea of any-granularity ranking, which leverages multi-vector embeddings to rank at varying levels of granularity while maintaining encoding at a single (coarser) level of granularity. We propose a multi-granular contrastive loss for training multi-vector approaches, and validate its utility with both sentences and propositions as ranking units. Finally, we demonstrate the application of proposition-level ranking to post-hoc citation addition in retrieval-augmented generation, surpassing the performance of prompt-driven citation generation. |
| title | AGRaME: Any-Granularity Ranking with Multi-Vector Embeddings |
| topic | Computation and Language Information Retrieval |
| url | https://arxiv.org/abs/2405.15028 |