A Representation Sharpening Framework for Zero Shot Dense Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ashok, Dhananjay, Nair, Suraj, Al-Darabsah, Mutasem, Teo, Choon Hui, Agarwal, Tarun, May, Jonathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918191825944576
author Ashok, Dhananjay
Nair, Suraj
Al-Darabsah, Mutasem
Teo, Choon Hui
Agarwal, Tarun
May, Jonathan
author_facet Ashok, Dhananjay
Nair, Suraj
Al-Darabsah, Mutasem
Teo, Choon Hui
Agarwal, Tarun
May, Jonathan
contents Zero-shot dense retrieval is a challenging setting where a document corpus is provided without relevant queries, necessitating a reliance on pretrained dense retrievers (DRs). However, since these DRs are not trained on the target corpus, they struggle to represent semantic differences between similar documents. To address this failing, we introduce a training-free representation sharpening framework that augments a document's representation with information that helps differentiate it from similar documents in the corpus. On over twenty datasets spanning multiple languages, the representation sharpening framework proves consistently superior to traditional retrieval, setting a new state-of-the-art on the BRIGHT benchmark. We show that representation sharpening is compatible with prior approaches to zero-shot dense retrieval and consistently improves their performance. Finally, we address the performance-cost tradeoff presented by our framework and devise an indexing-time approximation that preserves the majority of our performance gains over traditional retrieval, yet suffers no additional inference-time cost.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Representation Sharpening Framework for Zero Shot Dense Retrieval
Ashok, Dhananjay
Nair, Suraj
Al-Darabsah, Mutasem
Teo, Choon Hui
Agarwal, Tarun
May, Jonathan
Information Retrieval
Computation and Language
Zero-shot dense retrieval is a challenging setting where a document corpus is provided without relevant queries, necessitating a reliance on pretrained dense retrievers (DRs). However, since these DRs are not trained on the target corpus, they struggle to represent semantic differences between similar documents. To address this failing, we introduce a training-free representation sharpening framework that augments a document's representation with information that helps differentiate it from similar documents in the corpus. On over twenty datasets spanning multiple languages, the representation sharpening framework proves consistently superior to traditional retrieval, setting a new state-of-the-art on the BRIGHT benchmark. We show that representation sharpening is compatible with prior approaches to zero-shot dense retrieval and consistently improves their performance. Finally, we address the performance-cost tradeoff presented by our framework and devise an indexing-time approximation that preserves the majority of our performance gains over traditional retrieval, yet suffers no additional inference-time cost.
title A Representation Sharpening Framework for Zero Shot Dense Retrieval
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2511.05684