Corpus-Steered Query Expansion with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lei, Yibin, Cao, Yu, Zhou, Tianyi, Shen, Tao, Yates, Andrew
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929258410016768
author Lei, Yibin
Cao, Yu
Zhou, Tianyi
Shen, Tao
Yates, Andrew
author_facet Lei, Yibin
Cao, Yu
Zhou, Tianyi
Shen, Tao
Yates, Andrew
contents Recent studies demonstrate that query expansions generated by large language models (LLMs) can considerably enhance information retrieval systems by generating hypothetical documents that answer the queries as expansions. However, challenges arise from misalignments between the expansions and the retrieval corpus, resulting in issues like hallucinations and outdated information due to the limited intrinsic knowledge of LLMs. Inspired by Pseudo Relevance Feedback (PRF), we introduce Corpus-Steered Query Expansion (CSQE) to promote the incorporation of knowledge embedded within the corpus. CSQE utilizes the relevance assessing capability of LLMs to systematically identify pivotal sentences in the initially-retrieved documents. These corpus-originated texts are subsequently used to expand the query together with LLM-knowledge empowered expansions, improving the relevance prediction between the query and the target documents. Extensive experiments reveal that CSQE exhibits strong performance without necessitating any training, especially with queries for which LLMs lack knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18031
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Corpus-Steered Query Expansion with Large Language Models
Lei, Yibin
Cao, Yu
Zhou, Tianyi
Shen, Tao
Yates, Andrew
Information Retrieval
Computation and Language
Recent studies demonstrate that query expansions generated by large language models (LLMs) can considerably enhance information retrieval systems by generating hypothetical documents that answer the queries as expansions. However, challenges arise from misalignments between the expansions and the retrieval corpus, resulting in issues like hallucinations and outdated information due to the limited intrinsic knowledge of LLMs. Inspired by Pseudo Relevance Feedback (PRF), we introduce Corpus-Steered Query Expansion (CSQE) to promote the incorporation of knowledge embedded within the corpus. CSQE utilizes the relevance assessing capability of LLMs to systematically identify pivotal sentences in the initially-retrieved documents. These corpus-originated texts are subsequently used to expand the query together with LLM-knowledge empowered expansions, improving the relevance prediction between the query and the target documents. Extensive experiments reveal that CSQE exhibits strong performance without necessitating any training, especially with queries for which LLMs lack knowledge.
title Corpus-Steered Query Expansion with Large Language Models
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2402.18031