Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Weiwei, Yan, Lingyong, Ma, Xinyu, Wang, Shuaiqiang, Ren, Pengjie, Chen, Zhumin, Yin, Dawei, Ren, Zhaochun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917880212226048
author Sun, Weiwei
Yan, Lingyong
Ma, Xinyu
Wang, Shuaiqiang
Ren, Pengjie
Chen, Zhumin
Yin, Dawei
Ren, Zhaochun
author_facet Sun, Weiwei
Yan, Lingyong
Ma, Xinyu
Wang, Shuaiqiang
Ren, Pengjie
Chen, Zhumin
Yin, Dawei
Ren, Zhaochun
contents Large Language Models (LLMs) have demonstrated remarkable zero-shot generalization across various language-related tasks, including search engines. However, existing work utilizes the generative ability of LLMs for Information Retrieval (IR) rather than direct passage ranking. The discrepancy between the pre-training objectives of LLMs and the ranking objective poses another challenge. In this paper, we first investigate generative LLMs such as ChatGPT and GPT-4 for relevance ranking in IR. Surprisingly, our experiments reveal that properly instructed LLMs can deliver competitive, even superior results to state-of-the-art supervised methods on popular IR benchmarks. Furthermore, to address concerns about data contamination of LLMs, we collect a new test set called NovelEval, based on the latest knowledge and aiming to verify the model's ability to rank unknown knowledge. Finally, to improve efficiency in real-world applications, we delve into the potential for distilling the ranking capabilities of ChatGPT into small specialized models using a permutation distillation scheme. Our evaluation results turn out that a distilled 440M model outperforms a 3B supervised model on the BEIR benchmark. The code to reproduce our results is available at www.github.com/sunnweiwei/RankGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2304_09542
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents
Sun, Weiwei
Yan, Lingyong
Ma, Xinyu
Wang, Shuaiqiang
Ren, Pengjie
Chen, Zhumin
Yin, Dawei
Ren, Zhaochun
Computation and Language
Information Retrieval
Large Language Models (LLMs) have demonstrated remarkable zero-shot generalization across various language-related tasks, including search engines. However, existing work utilizes the generative ability of LLMs for Information Retrieval (IR) rather than direct passage ranking. The discrepancy between the pre-training objectives of LLMs and the ranking objective poses another challenge. In this paper, we first investigate generative LLMs such as ChatGPT and GPT-4 for relevance ranking in IR. Surprisingly, our experiments reveal that properly instructed LLMs can deliver competitive, even superior results to state-of-the-art supervised methods on popular IR benchmarks. Furthermore, to address concerns about data contamination of LLMs, we collect a new test set called NovelEval, based on the latest knowledge and aiming to verify the model's ability to rank unknown knowledge. Finally, to improve efficiency in real-world applications, we delve into the potential for distilling the ranking capabilities of ChatGPT into small specialized models using a permutation distillation scheme. Our evaluation results turn out that a distilled 440M model outperforms a 3B supervised model on the BEIR benchmark. The code to reproduce our results is available at www.github.com/sunnweiwei/RankGPT.
title Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2304.09542