How do LLMs Compute Verbal Confidence

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kumaran, Dharshan, Conmy, Arthur, Barbero, Federico, Osindero, Simon, Patraucean, Viorica, Veličković, Petar
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913145967083520
author Kumaran, Dharshan
Conmy, Arthur
Barbero, Federico
Osindero, Simon
Patraucean, Viorica
Veličković, Petar
author_facet Kumaran, Dharshan
Conmy, Arthur
Barbero, Federico
Osindero, Simon
Patraucean, Viorica
Veličković, Petar
contents Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs internally generate such scores remains unknown. We address two questions: first, when confidence is computed -- just-in-time when requested, or automatically during answer generation and cached for later retrieval; and second, what verbal confidence represents -- token log-probabilities, or a richer evaluation of answer quality? Focusing on Gemma 3 27B (across TriviaQA, BigMath, and MMLU), Qwen 2.5 7B, and the reasoning model Magistral Small 24B, we provide convergent evidence for cached retrieval. Activation steering, patching, noising, and swap experiments reveal that confidence representations emerge at answer-adjacent positions before appearing at the verbalization site. Attention blocking pinpoints the information flow: confidence is gathered from answer tokens, cached at the first post-answer position, then retrieved for output. Critically, linear probing and variance partitioning reveal that these cached representations explain substantial variance in verbal confidence beyond token log-probabilities, suggesting a richer answer-quality evaluation rather than a simple fluency readout. These findings demonstrate that verbal confidence reflects automatic, sophisticated self-evaluation -- not post-hoc reconstruction -- with implications for understanding metacognition in LLMs and improving calibration.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17839
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How do LLMs Compute Verbal Confidence
Kumaran, Dharshan
Conmy, Arthur
Barbero, Federico
Osindero, Simon
Patraucean, Viorica
Veličković, Petar
Computation and Language
Artificial Intelligence
Machine Learning
Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from black-box models. However, how LLMs internally generate such scores remains unknown. We address two questions: first, when confidence is computed -- just-in-time when requested, or automatically during answer generation and cached for later retrieval; and second, what verbal confidence represents -- token log-probabilities, or a richer evaluation of answer quality? Focusing on Gemma 3 27B (across TriviaQA, BigMath, and MMLU), Qwen 2.5 7B, and the reasoning model Magistral Small 24B, we provide convergent evidence for cached retrieval. Activation steering, patching, noising, and swap experiments reveal that confidence representations emerge at answer-adjacent positions before appearing at the verbalization site. Attention blocking pinpoints the information flow: confidence is gathered from answer tokens, cached at the first post-answer position, then retrieved for output. Critically, linear probing and variance partitioning reveal that these cached representations explain substantial variance in verbal confidence beyond token log-probabilities, suggesting a richer answer-quality evaluation rather than a simple fluency readout. These findings demonstrate that verbal confidence reflects automatic, sophisticated self-evaluation -- not post-hoc reconstruction -- with implications for understanding metacognition in LLMs and improving calibration.
title How do LLMs Compute Verbal Confidence
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.17839