Fact-Level Confidence Calibration and Self-Correction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Yige, Xu, Bingbing, Tan, Hexiang, Sun, Fei, Xiao, Teng, Li, Wei, Shen, Huawei, Cheng, Xueqi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929599169953792
author Yuan, Yige
Xu, Bingbing
Tan, Hexiang
Sun, Fei
Xiao, Teng
Li, Wei
Shen, Huawei
Cheng, Xueqi
author_facet Yuan, Yige
Xu, Bingbing
Tan, Hexiang
Sun, Fei
Xiao, Teng
Li, Wei
Shen, Huawei
Cheng, Xueqi
contents Confidence calibration in LLMs, i.e., aligning their self-assessed confidence with the actual accuracy of their responses, enabling them to self-evaluate the correctness of their outputs. However, current calibration methods for LLMs typically estimate two scalars to represent overall response confidence and correctness, which is inadequate for long-form generation where the response includes multiple atomic facts and may be partially confident and correct. These methods also overlook the relevance of each fact to the query. To address these challenges, we propose a Fact-Level Calibration framework that operates at a finer granularity, calibrating confidence to relevance-weighted correctness at the fact level. Furthermore, comprehensive analysis under the framework inspired the development of Confidence-Guided Fact-level Self-Correction ($\textbf{ConFix}$), which uses high-confidence facts within a response as additional knowledge to improve low-confidence ones. Extensive experiments across four datasets and six models demonstrate that ConFix effectively mitigates hallucinations without requiring external knowledge sources such as retrieval systems.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13343
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fact-Level Confidence Calibration and Self-Correction
Yuan, Yige
Xu, Bingbing
Tan, Hexiang
Sun, Fei
Xiao, Teng
Li, Wei
Shen, Huawei
Cheng, Xueqi
Computation and Language
Artificial Intelligence
Confidence calibration in LLMs, i.e., aligning their self-assessed confidence with the actual accuracy of their responses, enabling them to self-evaluate the correctness of their outputs. However, current calibration methods for LLMs typically estimate two scalars to represent overall response confidence and correctness, which is inadequate for long-form generation where the response includes multiple atomic facts and may be partially confident and correct. These methods also overlook the relevance of each fact to the query. To address these challenges, we propose a Fact-Level Calibration framework that operates at a finer granularity, calibrating confidence to relevance-weighted correctness at the fact level. Furthermore, comprehensive analysis under the framework inspired the development of Confidence-Guided Fact-level Self-Correction ($\textbf{ConFix}$), which uses high-confidence facts within a response as additional knowledge to improve low-confidence ones. Extensive experiments across four datasets and six models demonstrate that ConFix effectively mitigates hallucinations without requiring external knowledge sources such as retrieval systems.
title Fact-Level Confidence Calibration and Self-Correction
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.13343