RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hu, Chi, Ge, Yuan, Ma, Xiangnan, Cao, Hang, Li, Qiang, Yang, Yonghua, Xiao, Tong, Zhu, Jingbo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929285801967616
author Hu, Chi
Ge, Yuan
Ma, Xiangnan
Cao, Hang
Li, Qiang
Yang, Yonghua
Xiao, Tong
Zhu, Jingbo
author_facet Hu, Chi
Ge, Yuan
Ma, Xiangnan
Cao, Hang
Li, Qiang
Yang, Yonghua
Xiao, Tong
Zhu, Jingbo
contents Large Language Models (LLMs) have achieved impressive performance across various reasoning tasks. However, even state-of-the-art LLMs such as ChatGPT are prone to logical errors during their reasoning processes. Existing solutions, such as deploying task-specific verifiers or voting over multiple reasoning paths, either require extensive human annotations or fail in scenarios with inconsistent responses. To address these challenges, we introduce RankPrompt, a new prompting method that enables LLMs to self-rank their responses without additional resources. RankPrompt breaks down the ranking problem into a series of comparisons among diverse responses, leveraging the inherent capabilities of LLMs to generate chains of comparison as contextual exemplars. Our experiments across 11 arithmetic and commonsense reasoning tasks show that RankPrompt significantly enhances the reasoning performance of ChatGPT and GPT-4, with improvements of up to 13%. Moreover, RankPrompt excels in LLM-based automatic evaluations for open-ended tasks, aligning with human judgments 74% of the time in the AlpacaEval dataset. It also exhibits robustness to variations in response order and consistency. Collectively, our results validate RankPrompt as an effective method for eliciting high-quality feedback from language models.
format Preprint
id arxiv_https___arxiv_org_abs_2403_12373
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners
Hu, Chi
Ge, Yuan
Ma, Xiangnan
Cao, Hang
Li, Qiang
Yang, Yonghua
Xiao, Tong
Zhu, Jingbo
Computation and Language
Large Language Models (LLMs) have achieved impressive performance across various reasoning tasks. However, even state-of-the-art LLMs such as ChatGPT are prone to logical errors during their reasoning processes. Existing solutions, such as deploying task-specific verifiers or voting over multiple reasoning paths, either require extensive human annotations or fail in scenarios with inconsistent responses. To address these challenges, we introduce RankPrompt, a new prompting method that enables LLMs to self-rank their responses without additional resources. RankPrompt breaks down the ranking problem into a series of comparisons among diverse responses, leveraging the inherent capabilities of LLMs to generate chains of comparison as contextual exemplars. Our experiments across 11 arithmetic and commonsense reasoning tasks show that RankPrompt significantly enhances the reasoning performance of ChatGPT and GPT-4, with improvements of up to 13%. Moreover, RankPrompt excels in LLM-based automatic evaluations for open-ended tasks, aligning with human judgments 74% of the time in the AlpacaEval dataset. It also exhibits robustness to variations in response order and consistency. Collectively, our results validate RankPrompt as an effective method for eliciting high-quality feedback from language models.
title RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners
topic Computation and Language
url https://arxiv.org/abs/2403.12373