Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Daynauth, Roland, Clarke, Christopher, Flautner, Krisztian, Tang, Lingjia, Mars, Jason
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917925587255296
author Daynauth, Roland
Clarke, Christopher
Flautner, Krisztian
Tang, Lingjia
Mars, Jason
author_facet Daynauth, Roland
Clarke, Christopher
Flautner, Krisztian
Tang, Lingjia
Mars, Jason
contents Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a predefined criterion. By collecting these comparisons, a ranking can be constructed using methods such as Elo. However, applying these algorithms as constructed in the context of LLM evaluation introduces several challenges. In this paper, we explore the effectiveness of ranking systems for head-to-head comparisons of LLMs. We formally define a set of fundamental principles for effective ranking and conduct a series of extensive evaluations on the robustness of several ranking algorithms in the context of LLMs. Our analysis uncovers key insights into the factors that affect ranking accuracy and efficiency, offering guidelines for selecting the most appropriate methods based on specific evaluation contexts and resource constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14483
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
Daynauth, Roland
Clarke, Christopher
Flautner, Krisztian
Tang, Lingjia
Mars, Jason
Computation and Language
Artificial Intelligence
Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a predefined criterion. By collecting these comparisons, a ranking can be constructed using methods such as Elo. However, applying these algorithms as constructed in the context of LLM evaluation introduces several challenges. In this paper, we explore the effectiveness of ranking systems for head-to-head comparisons of LLMs. We formally define a set of fundamental principles for effective ranking and conduct a series of extensive evaluations on the robustness of several ranking algorithms in the context of LLMs. Our analysis uncovers key insights into the factors that affect ranking accuracy and efficiency, offering guidelines for selecting the most appropriate methods based on specific evaluation contexts and resource constraints.
title Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.14483