Large Language Models for Information Retrieval: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yutao, Yuan, Huaying, Wang, Shuting, Liu, Jiongnan, Liu, Wenhan, Deng, Chenlong, Chen, Haonan, Liu, Zheng, Dou, Zhicheng, Wen, Ji-Rong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908543889702912
author Zhu, Yutao
Yuan, Huaying
Wang, Shuting
Liu, Jiongnan
Liu, Wenhan
Deng, Chenlong
Chen, Haonan
Liu, Zheng
Dou, Zhicheng
Wen, Ji-Rong
author_facet Zhu, Yutao
Yuan, Huaying
Wang, Shuting
Liu, Jiongnan
Liu, Wenhan
Deng, Chenlong
Chen, Haonan
Liu, Zheng
Dou, Zhicheng
Wen, Ji-Rong
contents As a primary means of information acquisition, information retrieval (IR) systems, such as search engines, have integrated themselves into our daily lives. These systems also serve as components of dialogue, question-answering, and recommender systems. The trajectory of IR has evolved dynamically from its origins in term-based methods to its integration with advanced neural models. While the neural models excel at capturing complex contextual signals and semantic nuances, thereby reshaping the IR landscape, they still face challenges such as data scarcity, interpretability, and the generation of contextually plausible yet potentially inaccurate responses. This evolution requires a combination of both traditional methods (such as term-based sparse retrieval methods with rapid response) and modern neural architectures (such as language models with powerful language understanding capacity). Meanwhile, the emergence of large language models (LLMs), typified by ChatGPT and GPT-4, has revolutionized natural language processing due to their remarkable language understanding, generation, generalization, and reasoning abilities. Consequently, recent research has sought to leverage LLMs to improve IR systems. Given the rapid evolution of this research trajectory, it is necessary to consolidate existing methodologies and provide nuanced insights through a comprehensive overview. In this survey, we delve into the confluence of LLMs and IR systems, including crucial aspects such as query rewriters, retrievers, rerankers, and readers. Additionally, we explore promising directions, such as search agents, within this expanding field.
format Preprint
id arxiv_https___arxiv_org_abs_2308_07107
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Large Language Models for Information Retrieval: A Survey
Zhu, Yutao
Yuan, Huaying
Wang, Shuting
Liu, Jiongnan
Liu, Wenhan
Deng, Chenlong
Chen, Haonan
Liu, Zheng
Dou, Zhicheng
Wen, Ji-Rong
Computation and Language
Information Retrieval
As a primary means of information acquisition, information retrieval (IR) systems, such as search engines, have integrated themselves into our daily lives. These systems also serve as components of dialogue, question-answering, and recommender systems. The trajectory of IR has evolved dynamically from its origins in term-based methods to its integration with advanced neural models. While the neural models excel at capturing complex contextual signals and semantic nuances, thereby reshaping the IR landscape, they still face challenges such as data scarcity, interpretability, and the generation of contextually plausible yet potentially inaccurate responses. This evolution requires a combination of both traditional methods (such as term-based sparse retrieval methods with rapid response) and modern neural architectures (such as language models with powerful language understanding capacity). Meanwhile, the emergence of large language models (LLMs), typified by ChatGPT and GPT-4, has revolutionized natural language processing due to their remarkable language understanding, generation, generalization, and reasoning abilities. Consequently, recent research has sought to leverage LLMs to improve IR systems. Given the rapid evolution of this research trajectory, it is necessary to consolidate existing methodologies and provide nuanced insights through a comprehensive overview. In this survey, we delve into the confluence of LLMs and IR systems, including crucial aspects such as query rewriters, retrievers, rerankers, and readers. Additionally, we explore promising directions, such as search agents, within this expanding field.
title Large Language Models for Information Retrieval: A Survey
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2308.07107