Black-box Uncertainty Quantification Method for LLM-as-a-Judge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910651469791232 |
|---|---|
| author | Wagner, Nico Desmond, Michael Nair, Rahul Ashktorab, Zahra Daly, Elizabeth M. Pan, Qian Cooper, Martín Santillán Johnson, James M. Geyer, Werner |
| author_facet | Wagner, Nico Desmond, Michael Nair, Rahul Ashktorab, Zahra Daly, Elizabeth M. Pan, Qian Cooper, Martín Santillán Johnson, James M. Geyer, Werner |
| contents | LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty of LLM-as-a-Judge evaluations. While uncertainty quantification has been well-studied in other domains, applying it effectively to LLMs poses unique challenges due to their complex decision-making capabilities and computational demands. In this paper, we introduce a novel method for quantifying uncertainty designed to enhance the trustworthiness of LLM-as-a-Judge evaluations. The method quantifies uncertainty by analyzing the relationships between generated assessments and possible ratings. By cross-evaluating these relationships and constructing a confusion matrix based on token probabilities, the method derives labels of high or low uncertainty. We evaluate our method across multiple benchmarks, demonstrating a strong correlation between the accuracy of LLM evaluations and the derived uncertainty scores. Our findings suggest that this method can significantly improve the reliability and consistency of LLM-as-a-Judge evaluations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_11594 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Black-box Uncertainty Quantification Method for LLM-as-a-Judge Wagner, Nico Desmond, Michael Nair, Rahul Ashktorab, Zahra Daly, Elizabeth M. Pan, Qian Cooper, Martín Santillán Johnson, James M. Geyer, Werner Machine Learning Artificial Intelligence LLM-as-a-Judge is a widely used method for evaluating the performance of Large Language Models (LLMs) across various tasks. We address the challenge of quantifying the uncertainty of LLM-as-a-Judge evaluations. While uncertainty quantification has been well-studied in other domains, applying it effectively to LLMs poses unique challenges due to their complex decision-making capabilities and computational demands. In this paper, we introduce a novel method for quantifying uncertainty designed to enhance the trustworthiness of LLM-as-a-Judge evaluations. The method quantifies uncertainty by analyzing the relationships between generated assessments and possible ratings. By cross-evaluating these relationships and constructing a confusion matrix based on token probabilities, the method derives labels of high or low uncertainty. We evaluate our method across multiple benchmarks, demonstrating a strong correlation between the accuracy of LLM evaluations and the derived uncertainty scores. Our findings suggest that this method can significantly improve the reliability and consistency of LLM-as-a-Judge evaluations. |
| title | Black-box Uncertainty Quantification Method for LLM-as-a-Judge |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2410.11594 |