Are LLM Decisions Faithful to Verbal Confidence?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiawei, Zhou, Yanfei, Devic, Siddartha, Fu, Deqing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915723772690432
author Wang, Jiawei
Zhou, Yanfei
Devic, Siddartha
Fu, Deqing
author_facet Wang, Jiawei
Zhou, Yanfei
Devic, Siddartha
Fu, Deqing
contents Large Language Models (LLMs) can produce surprisingly sophisticated estimates of their own uncertainty. However, it remains unclear to what extent this expressed confidence is tied to the reasoning, knowledge, or decision making of the model. To test this, we introduce $\textbf{RiskEval}$: a framework designed to evaluate whether models adjust their abstention policies in response to varying error penalties. Our evaluation of several frontier models reveals a critical dissociation: models are neither cost-aware when articulating their verbal confidence, nor strategically responsive when deciding whether to engage or abstain under high-penalty conditions. Even when extreme penalties render frequent abstention the mathematically optimal strategy, models almost never abstain, resulting in utility collapse. This indicates that calibrated verbal confidence scores may not be sufficient to create trustworthy and interpretable AI systems, as current models lack the strategic agency to convert uncertainty signals into optimal and risk-sensitive decisions.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07767
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Are LLM Decisions Faithful to Verbal Confidence?
Wang, Jiawei
Zhou, Yanfei
Devic, Siddartha
Fu, Deqing
Machine Learning
Computation and Language
Large Language Models (LLMs) can produce surprisingly sophisticated estimates of their own uncertainty. However, it remains unclear to what extent this expressed confidence is tied to the reasoning, knowledge, or decision making of the model. To test this, we introduce $\textbf{RiskEval}$: a framework designed to evaluate whether models adjust their abstention policies in response to varying error penalties. Our evaluation of several frontier models reveals a critical dissociation: models are neither cost-aware when articulating their verbal confidence, nor strategically responsive when deciding whether to engage or abstain under high-penalty conditions. Even when extreme penalties render frequent abstention the mathematically optimal strategy, models almost never abstain, resulting in utility collapse. This indicates that calibrated verbal confidence scores may not be sufficient to create trustworthy and interpretable AI systems, as current models lack the strategic agency to convert uncertainty signals into optimal and risk-sensitive decisions.
title Are LLM Decisions Faithful to Verbal Confidence?
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2601.07767