Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917925587255296 |
|---|---|
| author | Daynauth, Roland Clarke, Christopher Flautner, Krisztian Tang, Lingjia Mars, Jason |
| author_facet | Daynauth, Roland Clarke, Christopher Flautner, Krisztian Tang, Lingjia Mars, Jason |
| contents | Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a predefined criterion. By collecting these comparisons, a ranking can be constructed using methods such as Elo. However, applying these algorithms as constructed in the context of LLM evaluation introduces several challenges. In this paper, we explore the effectiveness of ranking systems for head-to-head comparisons of LLMs. We formally define a set of fundamental principles for effective ranking and conduct a series of extensive evaluations on the robustness of several ranking algorithms in the context of LLMs. Our analysis uncovers key insights into the factors that affect ranking accuracy and efficiency, offering guidelines for selecting the most appropriate methods based on specific evaluation contexts and resource constraints. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_14483 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat Daynauth, Roland Clarke, Christopher Flautner, Krisztian Tang, Lingjia Mars, Jason Computation and Language Artificial Intelligence Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a predefined criterion. By collecting these comparisons, a ranking can be constructed using methods such as Elo. However, applying these algorithms as constructed in the context of LLM evaluation introduces several challenges. In this paper, we explore the effectiveness of ranking systems for head-to-head comparisons of LLMs. We formally define a set of fundamental principles for effective ranking and conduct a series of extensive evaluations on the robustness of several ranking algorithms in the context of LLMs. Our analysis uncovers key insights into the factors that affect ranking accuracy and efficiency, offering guidelines for selecting the most appropriate methods based on specific evaluation contexts and resource constraints. |
| title | Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2411.14483 |