Closing the Confidence-Faithfulness Gap in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miao, Miranda Muqing, Ungar, Lyle
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908930052980736
author Miao, Miranda Muqing
Ungar, Lyle
author_facet Miao, Miranda Muqing
Ungar, Lyle
contents Large language models (LLMs) tend to verbalize confidence scores that are largely detached from their actual accuracy, yet the geometric relationship governing this behavior remain poorly understood. In this work, we present a mechanistic interpretability analysis of verbalized confidence, using linear probes and contrastive activation addition (CAA) steering to show that calibration and verbalized confidence signals are encoded linearly but are orthogonal to one another -- a finding consistent across three open-weight models and four datasets. Interestingly, when models are prompted to simultaneously reason through a problem and verbalize a confidence score, the reasoning process disrupts the verbalized confidence direction, exacerbating miscalibration. We term this the "Reasoning Contamination Effect." Leveraging this insight, we introduce a two-stage adaptive steering pipeline that reads the model's internal accuracy estimate and steers verbalized output to match it, substantially improving calibration alignment across all evaluated models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25052
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Closing the Confidence-Faithfulness Gap in Large Language Models
Miao, Miranda Muqing
Ungar, Lyle
Computation and Language
Artificial Intelligence
Large language models (LLMs) tend to verbalize confidence scores that are largely detached from their actual accuracy, yet the geometric relationship governing this behavior remain poorly understood. In this work, we present a mechanistic interpretability analysis of verbalized confidence, using linear probes and contrastive activation addition (CAA) steering to show that calibration and verbalized confidence signals are encoded linearly but are orthogonal to one another -- a finding consistent across three open-weight models and four datasets. Interestingly, when models are prompted to simultaneously reason through a problem and verbalize a confidence score, the reasoning process disrupts the verbalized confidence direction, exacerbating miscalibration. We term this the "Reasoning Contamination Effect." Leveraging this insight, we introduce a two-stage adaptive steering pipeline that reads the model's internal accuracy estimate and steers verbalized output to match it, substantially improving calibration alignment across all evaluated models.
title Closing the Confidence-Faithfulness Gap in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.25052