An Investigation of Prompt Variations for Zero-shot LLM-based Rankers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Shuoqi, Zhuang, Shengyao, Wang, Shuai, Zuccon, Guido
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915408660922368
author Sun, Shuoqi
Zhuang, Shengyao
Wang, Shuai
Zuccon, Guido
author_facet Sun, Shuoqi
Zhuang, Shengyao
Wang, Shuai
Zuccon, Guido
contents We provide a systematic understanding of the impact of specific components and wordings used in prompts on the effectiveness of rankers based on zero-shot Large Language Models (LLMs). Several zero-shot ranking methods based on LLMs have recently been proposed. Among many aspects, methods differ across (1) the ranking algorithm they implement, e.g., pointwise vs. listwise, (2) the backbone LLMs used, e.g., GPT3.5 vs. FLAN-T5, (3) the components and wording used in prompts, e.g., the use or not of role-definition (role-playing) and the actual words used to express this. It is currently unclear whether performance differences are due to the underlying ranking algorithm, or because of spurious factors such as better choice of words used in prompts. This confusion risks to undermine future research. Through our large-scale experimentation and analysis, we find that ranking algorithms do contribute to differences between methods for zero-shot LLM ranking. However, so do the LLM backbones -- but even more importantly, the choice of prompt components and wordings affect the ranking. In fact, in our experiments, we find that, at times, these latter elements have more impact on the ranker's effectiveness than the actual ranking algorithms, and that differences among ranking methods become more blurred when prompt variations are considered.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14117
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Investigation of Prompt Variations for Zero-shot LLM-based Rankers
Sun, Shuoqi
Zhuang, Shengyao
Wang, Shuai
Zuccon, Guido
Information Retrieval
Computation and Language
We provide a systematic understanding of the impact of specific components and wordings used in prompts on the effectiveness of rankers based on zero-shot Large Language Models (LLMs). Several zero-shot ranking methods based on LLMs have recently been proposed. Among many aspects, methods differ across (1) the ranking algorithm they implement, e.g., pointwise vs. listwise, (2) the backbone LLMs used, e.g., GPT3.5 vs. FLAN-T5, (3) the components and wording used in prompts, e.g., the use or not of role-definition (role-playing) and the actual words used to express this. It is currently unclear whether performance differences are due to the underlying ranking algorithm, or because of spurious factors such as better choice of words used in prompts. This confusion risks to undermine future research. Through our large-scale experimentation and analysis, we find that ranking algorithms do contribute to differences between methods for zero-shot LLM ranking. However, so do the LLM backbones -- but even more importantly, the choice of prompt components and wordings affect the ranking. In fact, in our experiments, we find that, at times, these latter elements have more impact on the ranker's effectiveness than the actual ranking algorithms, and that differences among ranking methods become more blurred when prompt variations are considered.
title An Investigation of Prompt Variations for Zero-shot LLM-based Rankers
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2406.14117