Calibrating Large Language Models with Sample Consistency

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lyu, Qing, Shridhar, Kumar, Malaviya, Chaitanya, Zhang, Li, Elazar, Yanai, Tandon, Niket, Apidianaki, Marianna, Sachan, Mrinmaya, Callison-Burch, Chris
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914342542245888
author Lyu, Qing
Shridhar, Kumar
Malaviya, Chaitanya
Zhang, Li
Elazar, Yanai
Tandon, Niket
Apidianaki, Marianna
Sachan, Mrinmaya
Callison-Burch, Chris
author_facet Lyu, Qing
Shridhar, Kumar
Malaviya, Chaitanya
Zhang, Li
Elazar, Yanai
Tandon, Niket
Apidianaki, Marianna
Sachan, Mrinmaya
Callison-Burch, Chris
contents Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we explore the potential of deriving confidence from the distribution of multiple randomly sampled model generations, via three measures of consistency. We perform an extensive evaluation across various open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency have the potential to enhance model performance. Finally, we offer practical guidance on choosing suitable consistency metrics for calibration, tailored to the characteristics of various LMs.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13904
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Calibrating Large Language Models with Sample Consistency
Lyu, Qing
Shridhar, Kumar
Malaviya, Chaitanya
Zhang, Li
Elazar, Yanai
Tandon, Niket
Apidianaki, Marianna
Sachan, Mrinmaya
Callison-Burch, Chris
Computation and Language
Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we explore the potential of deriving confidence from the distribution of multiple randomly sampled model generations, via three measures of consistency. We perform an extensive evaluation across various open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency have the potential to enhance model performance. Finally, we offer practical guidance on choosing suitable consistency metrics for calibration, tailored to the characteristics of various LMs.
title Calibrating Large Language Models with Sample Consistency
topic Computation and Language
url https://arxiv.org/abs/2402.13904