Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yoon, Yejun, Jung, Jaeyoon, Yoon, Seunghyun, Park, Kunwoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908392202698752
author Yoon, Yejun
Jung, Jaeyoon
Yoon, Seunghyun
Park, Kunwoo
author_facet Yoon, Yejun
Jung, Jaeyoon
Yoon, Seunghyun
Park, Kunwoo
contents Query expansion methods powered by large language models (LLMs) have demonstrated effectiveness in zero-shot retrieval tasks. These methods assume that LLMs can generate hypothetical documents that, when incorporated into a query vector, enhance the retrieval of real evidence. However, we challenge this assumption by investigating whether knowledge leakage in benchmarks contributes to the observed performance gains. Using fact verification as a testbed, we analyze whether the generated documents contain information entailed by ground-truth evidence and assess their impact on performance. Our findings indicate that, on average, performance improvements consistently occurred for claims whose generated documents included sentences entailed by gold evidence. This suggests that knowledge leakage may be present in fact-verification benchmarks, potentially inflating the perceived performance of LLM-based query expansion methods.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14175
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
Yoon, Yejun
Jung, Jaeyoon
Yoon, Seunghyun
Park, Kunwoo
Computation and Language
Information Retrieval
Query expansion methods powered by large language models (LLMs) have demonstrated effectiveness in zero-shot retrieval tasks. These methods assume that LLMs can generate hypothetical documents that, when incorporated into a query vector, enhance the retrieval of real evidence. However, we challenge this assumption by investigating whether knowledge leakage in benchmarks contributes to the observed performance gains. Using fact verification as a testbed, we analyze whether the generated documents contain information entailed by ground-truth evidence and assess their impact on performance. Our findings indicate that, on average, performance improvements consistently occurred for claims whose generated documents included sentences entailed by gold evidence. This suggests that knowledge leakage may be present in fact-verification benchmarks, potentially inflating the perceived performance of LLM-based query expansion methods.
title Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2504.14175