Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917501655318528 |
|---|---|
| author | Lu, Yuyin Liang, Ziran Rao, Yanghui Fan, Wenqi Wang, Fu Lee Li, Qing |
| author_facet | Lu, Yuyin Liang, Ziran Rao, Yanghui Fan, Wenqi Wang, Fu Lee Li, Qing |
| contents | Reliable reasoning in Large Language Models (LLMs) is challenged by their propensity for hallucination. While augmenting LLMs with Knowledge Graphs (KGs) improves factual accuracy, existing KG-augmented methods fail to quantify epistemic uncertainty in both the retrieved evidence and LLMs' reasoning. To bridge this gap, we introduce DoublyCal, a framework built on a novel double-calibration principle. DoublyCal employs a lightweight proxy model to first generate KG evidence alongside a calibrated evidence confidence. This calibrated supporting evidence then guides a black-box LLM, yielding final predictions that are not only more accurate but also well-calibrated, with confidence scores traceable to the uncertainty of the supporting evidence. Experiments on knowledge-intensive benchmarks show that DoublyCal significantly improves both the accuracy and confidence calibration of black-box LLMs while maintaining low token cost. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_11956 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence Lu, Yuyin Liang, Ziran Rao, Yanghui Fan, Wenqi Wang, Fu Lee Li, Qing Computation and Language Artificial Intelligence Reliable reasoning in Large Language Models (LLMs) is challenged by their propensity for hallucination. While augmenting LLMs with Knowledge Graphs (KGs) improves factual accuracy, existing KG-augmented methods fail to quantify epistemic uncertainty in both the retrieved evidence and LLMs' reasoning. To bridge this gap, we introduce DoublyCal, a framework built on a novel double-calibration principle. DoublyCal employs a lightweight proxy model to first generate KG evidence alongside a calibrated evidence confidence. This calibrated supporting evidence then guides a black-box LLM, yielding final predictions that are not only more accurate but also well-calibrated, with confidence scores traceable to the uncertainty of the supporting evidence. Experiments on knowledge-intensive benchmarks show that DoublyCal significantly improves both the accuracy and confidence calibration of black-box LLMs while maintaining low token cost. |
| title | Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2601.11956 |