Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Zebin, Han, Yi, Fang, Ethan X., Wang, Lan, Lu, Junwei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912225144340480
author Wang, Zebin
Han, Yi
Fang, Ethan X.
Wang, Lan
Lu, Junwei
author_facet Wang, Zebin
Han, Yi
Fang, Ethan X.
Wang, Lan
Lu, Junwei
contents We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs has proven to be an effective tool to improve alignment based on the best-of-$N$ policy. In this paper, we propose a new inferential framework for hypothesis testing among the ranking for language models. Our framework is based on a nonparametric contextual ranking framework designed to assess large language models' domain-specific expertise, leveraging nonparametric scoring methods to account for their sensitivity to the prompts. To characterize the combinatorial complexity of the ranking, we introduce a novel concept of confidence diagram, which leverages a Hasse diagram to represent the entire confidence set of rankings by a single directed graph. We show the validity of the proposed confidence diagram by advancing the Gaussian multiplier bootstrap theory to accommodate the supremum of independent empirical processes that are not necessarily identically distributed. Extensive numerical experiments conducted on both synthetic and real data demonstrate that our approach offers valuable insight into the evaluation for the performance of different LLMs across various medical domains.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05506
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation
Wang, Zebin
Han, Yi
Fang, Ethan X.
Wang, Lan
Lu, Junwei
Machine Learning
Methodology
We consider the inference for the ranking of large language models (LLMs). Alignment arises as a significant challenge to mitigate hallucinations in the use of LLMs. Ranking LLMs has proven to be an effective tool to improve alignment based on the best-of-$N$ policy. In this paper, we propose a new inferential framework for hypothesis testing among the ranking for language models. Our framework is based on a nonparametric contextual ranking framework designed to assess large language models' domain-specific expertise, leveraging nonparametric scoring methods to account for their sensitivity to the prompts. To characterize the combinatorial complexity of the ranking, we introduce a novel concept of confidence diagram, which leverages a Hasse diagram to represent the entire confidence set of rankings by a single directed graph. We show the validity of the proposed confidence diagram by advancing the Gaussian multiplier bootstrap theory to accommodate the supremum of independent empirical processes that are not necessarily identically distributed. Extensive numerical experiments conducted on both synthetic and real data demonstrate that our approach offers valuable insight into the evaluation for the performance of different LLMs across various medical domains.
title Confidence Diagram of Nonparametric Ranking for Uncertainty Assessment in Large Language Models Evaluation
topic Machine Learning
Methodology
url https://arxiv.org/abs/2412.05506