Black-box Uncertainty Quantification Method for LLM-as-a-Judge

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wagner, Nico, Desmond, Michael, Nair, Rahul, Ashktorab, Zahra, Daly, Elizabeth M., Pan, Qian, Cooper, Martín Santillán, Johnson, James M., Geyer, Werner
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910651469791232
author Wagner, Nico
Desmond, Michael
Nair, Rahul
Ashktorab, Zahra
Daly, Elizabeth M.
Pan, Qian
Cooper, Martín Santillán
Johnson, James M.
Geyer, Werner
author_facet Wagner, Nico
Desmond, Michael
Nair, Rahul
Ashktorab, Zahra
Daly, Elizabeth M.
Pan, Qian
Cooper, Martín Santillán
Johnson, James M.
Geyer, Werner
contents LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty of LLM-as-a-Judge evaluations. While uncertainty quantification has been well-studied in other domains, applying it effectively to LLMs poses unique challenges due to their complex decision-making capabilities and computational demands. In this paper, we introduce a novel method for quantifying uncertainty designed to enhance the trustworthiness of LLM-as-a-Judge evaluations. The method quantifies uncertainty by analyzing the relationships between generated assessments and possible ratings. By cross-evaluating these relationships and constructing a confusion matrix based on token probabilities, the method derives labels of high or low uncertainty. We evaluate our method across multiple benchmarks, demonstrating a strong correlation between the accuracy of LLM evaluations and the derived uncertainty scores. Our findings suggest that this method can significantly improve the reliability and consistency of LLM-as-a-Judge evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11594
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Black-box Uncertainty Quantification Method for LLM-as-a-Judge
Wagner, Nico
Desmond, Michael
Nair, Rahul
Ashktorab, Zahra
Daly, Elizabeth M.
Pan, Qian
Cooper, Martín Santillán
Johnson, James M.
Geyer, Werner
Machine Learning
Artificial Intelligence
LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty of LLM-as-a-Judge evaluations. While uncertainty quantification has been well-studied in other domains, applying it effectively to LLMs poses unique challenges due to their complex decision-making capabilities and computational demands. In this paper, we introduce a novel method for quantifying uncertainty designed to enhance the trustworthiness of LLM-as-a-Judge evaluations. The method quantifies uncertainty by analyzing the relationships between generated assessments and possible ratings. By cross-evaluating these relationships and constructing a confusion matrix based on token probabilities, the method derives labels of high or low uncertainty. We evaluate our method across multiple benchmarks, demonstrating a strong correlation between the accuracy of LLM evaluations and the derived uncertainty scores. Our findings suggest that this method can significantly improve the reliability and consistency of LLM-as-a-Judge evaluations.
title Black-box Uncertainty Quantification Method for LLM-as-a-Judge
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.11594