Calibrating the Confidence of Large Language Models by Eliciting Fidelity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Mozhi, Huang, Mianqiu, Shi, Rundong, Guo, Linsen, Peng, Chong, Yan, Peng, Zhou, Yaqian, Qiu, Xipeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909341113647104
author Zhang, Mozhi
Huang, Mianqiu
Shi, Rundong
Guo, Linsen
Peng, Chong
Yan, Peng
Zhou, Yaqian
Qiu, Xipeng
author_facet Zhang, Mozhi
Huang, Mianqiu
Shi, Rundong
Guo, Linsen
Peng, Chong
Yan, Peng
Zhou, Yaqian
Qiu, Xipeng
contents Large language models optimized with techniques like RLHF have achieved good alignment in being helpful and harmless. However, post-alignment, these language models often exhibit overconfidence, where the expressed confidence does not accurately calibrate with their correctness rate. In this paper, we decompose the language model confidence into the \textit{Uncertainty} about the question and the \textit{Fidelity} to the answer generated by language models. Then, we propose a plug-and-play method to estimate the confidence of language models. Our method has shown good calibration performance by conducting experiments with 6 RLHF-LMs on four MCQA datasets. Moreover, we propose two novel metrics, IPR and CE, to evaluate the calibration of the model, and we have conducted a detailed discussion on \textit{Truly Well-Calibrated Confidence}. Our method could serve as a strong baseline, and we hope that this work will provide some insights into the model confidence calibration.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02655
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Calibrating the Confidence of Large Language Models by Eliciting Fidelity
Zhang, Mozhi
Huang, Mianqiu
Shi, Rundong
Guo, Linsen
Peng, Chong
Yan, Peng
Zhou, Yaqian
Qiu, Xipeng
Computation and Language
Large language models optimized with techniques like RLHF have achieved good alignment in being helpful and harmless. However, post-alignment, these language models often exhibit overconfidence, where the expressed confidence does not accurately calibrate with their correctness rate. In this paper, we decompose the language model confidence into the \textit{Uncertainty} about the question and the \textit{Fidelity} to the answer generated by language models. Then, we propose a plug-and-play method to estimate the confidence of language models. Our method has shown good calibration performance by conducting experiments with 6 RLHF-LMs on four MCQA datasets. Moreover, we propose two novel metrics, IPR and CE, to evaluate the calibration of the model, and we have conducted a detailed discussion on \textit{Truly Well-Calibrated Confidence}. Our method could serve as a strong baseline, and we hope that this work will provide some insights into the model confidence calibration.
title Calibrating the Confidence of Large Language Models by Eliciting Fidelity
topic Computation and Language
url https://arxiv.org/abs/2404.02655