LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abe, Kenya, Takeoka, Kunihiro, Kato, Makoto P., Oyamada, Masafumi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916742924599296
author Abe, Kenya
Takeoka, Kunihiro
Kato, Makoto P.
Oyamada, Masafumi
author_facet Abe, Kenya
Takeoka, Kunihiro
Kato, Makoto P.
Oyamada, Masafumi
contents Query expansion (QE) enhances retrieval by incorporating relevant terms, with large language models (LLMs) offering an effective alternative to traditional rule-based and statistical methods. However, LLM-based QE suffers from a fundamental limitation: it often fails to generate relevant knowledge, degrading search performance. Prior studies have focused on hallucination, yet its underlying cause--LLM knowledge deficiencies--remains underexplored. This paper systematically examines two failure cases in LLM-based QE: (1) when the LLM lacks query knowledge, leading to incorrect expansions, and (2) when the query is ambiguous, causing biased refinements that narrow search coverage. We conduct controlled experiments across multiple datasets, evaluating the effects of knowledge and query ambiguity on retrieval performance using sparse and dense retrieval models. Our results reveal that LLM-based QE can significantly degrade the retrieval effectiveness when knowledge in the LLM is insufficient or query ambiguity is high. We introduce a framework for evaluating QE under these conditions, providing insights into the limitations of LLM-based retrieval augmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries
Abe, Kenya
Takeoka, Kunihiro
Kato, Makoto P.
Oyamada, Masafumi
Information Retrieval
H.3.3
Query expansion (QE) enhances retrieval by incorporating relevant terms, with large language models (LLMs) offering an effective alternative to traditional rule-based and statistical methods. However, LLM-based QE suffers from a fundamental limitation: it often fails to generate relevant knowledge, degrading search performance. Prior studies have focused on hallucination, yet its underlying cause--LLM knowledge deficiencies--remains underexplored. This paper systematically examines two failure cases in LLM-based QE: (1) when the LLM lacks query knowledge, leading to incorrect expansions, and (2) when the query is ambiguous, causing biased refinements that narrow search coverage. We conduct controlled experiments across multiple datasets, evaluating the effects of knowledge and query ambiguity on retrieval performance using sparse and dense retrieval models. Our results reveal that LLM-based QE can significantly degrade the retrieval effectiveness when knowledge in the LLM is insufficient or query ambiguity is high. We introduce a framework for evaluating QE under these conditions, providing insights into the limitations of LLM-based retrieval augmentation.
title LLM-based Query Expansion Fails for Unfamiliar and Ambiguous Queries
topic Information Retrieval
H.3.3
url https://arxiv.org/abs/2505.12694