Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yuyin, Liang, Ziran, Rao, Yanghui, Fan, Wenqi, Wang, Fu Lee, Li, Qing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917501655318528
author Lu, Yuyin
Liang, Ziran
Rao, Yanghui
Fan, Wenqi
Wang, Fu Lee
Li, Qing
author_facet Lu, Yuyin
Liang, Ziran
Rao, Yanghui
Fan, Wenqi
Wang, Fu Lee
Li, Qing
contents Reliable reasoning in Large Language Models (LLMs) is challenged by their propensity for hallucination. While augmenting LLMs with Knowledge Graphs (KGs) improves factual accuracy, existing KG-augmented methods fail to quantify epistemic uncertainty in both the retrieved evidence and LLMs' reasoning. To bridge this gap, we introduce DoublyCal, a framework built on a novel double-calibration principle. DoublyCal employs a lightweight proxy model to first generate KG evidence alongside a calibrated evidence confidence. This calibrated supporting evidence then guides a black-box LLM, yielding final predictions that are not only more accurate but also well-calibrated, with confidence scores traceable to the uncertainty of the supporting evidence. Experiments on knowledge-intensive benchmarks show that DoublyCal significantly improves both the accuracy and confidence calibration of black-box LLMs while maintaining low token cost.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11956
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
Lu, Yuyin
Liang, Ziran
Rao, Yanghui
Fan, Wenqi
Wang, Fu Lee
Li, Qing
Computation and Language
Artificial Intelligence
Reliable reasoning in Large Language Models (LLMs) is challenged by their propensity for hallucination. While augmenting LLMs with Knowledge Graphs (KGs) improves factual accuracy, existing KG-augmented methods fail to quantify epistemic uncertainty in both the retrieved evidence and LLMs' reasoning. To bridge this gap, we introduce DoublyCal, a framework built on a novel double-calibration principle. DoublyCal employs a lightweight proxy model to first generate KG evidence alongside a calibrated evidence confidence. This calibrated supporting evidence then guides a black-box LLM, yielding final predictions that are not only more accurate but also well-calibrated, with confidence scores traceable to the uncertainty of the supporting evidence. Experiments on knowledge-intensive benchmarks show that DoublyCal significantly improves both the accuracy and confidence calibration of black-box LLMs while maintaining low token cost.
title Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.11956