Zero-Shot Dense Retrieval with Embeddings from Relevance Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jedidi, Nour, Chuang, Yung-Sung, Shing, Leslie, Glass, James
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910673524490240
author Jedidi, Nour
Chuang, Yung-Sung
Shing, Leslie
Glass, James
author_facet Jedidi, Nour
Chuang, Yung-Sung
Shing, Leslie
Glass, James
contents Building effective dense retrieval systems remains difficult when relevance supervision is not available. Recent work has looked to overcome this challenge by using a Large Language Model (LLM) to generate hypothetical documents that can be used to find the closest real document. However, this approach relies solely on the LLM to have domain-specific knowledge relevant to the query, which may not be practical. Furthermore, generating hypothetical documents can be inefficient as it requires the LLM to generate a large number of tokens for each query. To address these challenges, we introduce Real Document Embeddings from Relevance Feedback (ReDE-RF). Inspired by relevance feedback, ReDE-RF proposes to re-frame hypothetical document generation as a relevance estimation task, using an LLM to select which documents should be used for nearest neighbor search. Through this re-framing, the LLM no longer needs domain-specific knowledge but only needs to judge what is relevant. Additionally, relevance estimation only requires the LLM to output a single token, thereby improving search latency. Our experiments show that ReDE-RF consistently surpasses state-of-the-art zero-shot dense retrieval methods across a wide range of low-resource retrieval datasets while also making significant improvements in latency per-query.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21242
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-Shot Dense Retrieval with Embeddings from Relevance Feedback
Jedidi, Nour
Chuang, Yung-Sung
Shing, Leslie
Glass, James
Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
Building effective dense retrieval systems remains difficult when relevance supervision is not available. Recent work has looked to overcome this challenge by using a Large Language Model (LLM) to generate hypothetical documents that can be used to find the closest real document. However, this approach relies solely on the LLM to have domain-specific knowledge relevant to the query, which may not be practical. Furthermore, generating hypothetical documents can be inefficient as it requires the LLM to generate a large number of tokens for each query. To address these challenges, we introduce Real Document Embeddings from Relevance Feedback (ReDE-RF). Inspired by relevance feedback, ReDE-RF proposes to re-frame hypothetical document generation as a relevance estimation task, using an LLM to select which documents should be used for nearest neighbor search. Through this re-framing, the LLM no longer needs domain-specific knowledge but only needs to judge what is relevant. Additionally, relevance estimation only requires the LLM to output a single token, thereby improving search latency. Our experiments show that ReDE-RF consistently surpasses state-of-the-art zero-shot dense retrieval methods across a wide range of low-resource retrieval datasets while also making significant improvements in latency per-query.
title Zero-Shot Dense Retrieval with Embeddings from Relevance Feedback
topic Information Retrieval
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.21242