A Representation Sharpening Framework for Zero Shot Dense Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918191825944576 |
|---|---|
| author | Ashok, Dhananjay Nair, Suraj Al-Darabsah, Mutasem Teo, Choon Hui Agarwal, Tarun May, Jonathan |
| author_facet | Ashok, Dhananjay Nair, Suraj Al-Darabsah, Mutasem Teo, Choon Hui Agarwal, Tarun May, Jonathan |
| contents | Zero-shot dense retrieval is a challenging setting where a document corpus is provided without relevant queries, necessitating a reliance on pretrained dense retrievers (DRs). However, since these DRs are not trained on the target corpus, they struggle to represent semantic differences between similar documents. To address this failing, we introduce a training-free representation sharpening framework that augments a document's representation with information that helps differentiate it from similar documents in the corpus. On over twenty datasets spanning multiple languages, the representation sharpening framework proves consistently superior to traditional retrieval, setting a new state-of-the-art on the BRIGHT benchmark. We show that representation sharpening is compatible with prior approaches to zero-shot dense retrieval and consistently improves their performance. Finally, we address the performance-cost tradeoff presented by our framework and devise an indexing-time approximation that preserves the majority of our performance gains over traditional retrieval, yet suffers no additional inference-time cost. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_05684 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Representation Sharpening Framework for Zero Shot Dense Retrieval Ashok, Dhananjay Nair, Suraj Al-Darabsah, Mutasem Teo, Choon Hui Agarwal, Tarun May, Jonathan Information Retrieval Computation and Language Zero-shot dense retrieval is a challenging setting where a document corpus is provided without relevant queries, necessitating a reliance on pretrained dense retrievers (DRs). However, since these DRs are not trained on the target corpus, they struggle to represent semantic differences between similar documents. To address this failing, we introduce a training-free representation sharpening framework that augments a document's representation with information that helps differentiate it from similar documents in the corpus. On over twenty datasets spanning multiple languages, the representation sharpening framework proves consistently superior to traditional retrieval, setting a new state-of-the-art on the BRIGHT benchmark. We show that representation sharpening is compatible with prior approaches to zero-shot dense retrieval and consistently improves their performance. Finally, we address the performance-cost tradeoff presented by our framework and devise an indexing-time approximation that preserves the majority of our performance gains over traditional retrieval, yet suffers no additional inference-time cost. |
| title | A Representation Sharpening Framework for Zero Shot Dense Retrieval |
| topic | Information Retrieval Computation and Language |
| url | https://arxiv.org/abs/2511.05684 |