Calibrating Large Language Models Using Their Generations Only

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ulmer, Dennis, Gubri, Martin, Lee, Hwaran, Yun, Sangdoo, Oh, Seong Joon
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913260452708352
author Ulmer, Dennis
Gubri, Martin
Lee, Hwaran
Yun, Sangdoo
Oh, Seong Joon
author_facet Ulmer, Dennis
Gubri, Martin
Lee, Hwaran
Yun, Sangdoo
Oh, Seong Joon
contents As large language models (LLMs) are increasingly deployed in user-facing applications, building trust and maintaining safety by accurately quantifying a model's confidence in its prediction becomes even more important. However, finding effective ways to calibrate LLMs - especially when the only interface to the models is their generated text - remains a challenge. We propose APRICOT (auxiliary prediction of confidence targets): A method to set confidence targets and train an additional model that predicts an LLM's confidence based on its textual input and output alone. This approach has several advantages: It is conceptually simple, does not require access to the target model beyond its output, does not interfere with the language generation, and has a multitude of potential usages, for instance by verbalizing the predicted confidence or adjusting the given answer based on the confidence. We show how our approach performs competitively in terms of calibration error for white-box and black-box LLMs on closed-book question-answering to detect incorrect LLM answers.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05973
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Calibrating Large Language Models Using Their Generations Only
Ulmer, Dennis
Gubri, Martin
Lee, Hwaran
Yun, Sangdoo
Oh, Seong Joon
Computation and Language
Artificial Intelligence
Machine Learning
As large language models (LLMs) are increasingly deployed in user-facing applications, building trust and maintaining safety by accurately quantifying a model's confidence in its prediction becomes even more important. However, finding effective ways to calibrate LLMs - especially when the only interface to the models is their generated text - remains a challenge. We propose APRICOT (auxiliary prediction of confidence targets): A method to set confidence targets and train an additional model that predicts an LLM's confidence based on its textual input and output alone. This approach has several advantages: It is conceptually simple, does not require access to the target model beyond its output, does not interfere with the language generation, and has a multitude of potential usages, for instance by verbalizing the predicted confidence or adjusting the given answer based on the confidence. We show how our approach performs competitively in terms of calibration error for white-box and black-box LLMs on closed-book question-answering to detect incorrect LLM answers.
title Calibrating Large Language Models Using Their Generations Only
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.05973