Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lewis-Lim, Samuel, Tan, Xingwei, Zhao, Zhixue, Aletras, Nikolaos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915716289003520
author Lewis-Lim, Samuel
Tan, Xingwei
Zhao, Zhixue
Aletras, Nikolaos
author_facet Lewis-Lim, Samuel
Tan, Xingwei
Zhao, Zhixue
Aletras, Nikolaos
contents Chain-of-thought (CoT) prompting is a common technique for improving the reasoning abilities of large language models (LLMs). However, extended reasoning is often unnecessary and substantially increases token usage. As such, a key question becomes how to optimally allocate compute to when reasoning is actually needed. We study this through confidence-gated CoT, where a model produces a direct answer and a confidence estimate to decide whether to invoke CoT. We present an evaluation framework together with the first systematic study of confidence signals for this decision. We evaluate four representative confidence measures and compare them with random gating and an oracle upper bound. Experiments across two model families and diverse reasoning tasks show that existing training-free confidence measures can reduce redundant reasoning. However, we also find that the utility of individual confidence measures is inconsistent across settings. Through our evaluation framework and analysis, our study provides practical guidance toward developing and evaluating models that selectively use CoT.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21007
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
Lewis-Lim, Samuel
Tan, Xingwei
Zhao, Zhixue
Aletras, Nikolaos
Computation and Language
Chain-of-thought (CoT) prompting is a common technique for improving the reasoning abilities of large language models (LLMs). However, extended reasoning is often unnecessary and substantially increases token usage. As such, a key question becomes how to optimally allocate compute to when reasoning is actually needed. We study this through confidence-gated CoT, where a model produces a direct answer and a confidence estimate to decide whether to invoke CoT. We present an evaluation framework together with the first systematic study of confidence signals for this decision. We evaluate four representative confidence measures and compare them with random gating and an oracle upper bound. Experiments across two model families and diverse reasoning tasks show that existing training-free confidence measures can reduce redundant reasoning. However, we also find that the utility of individual confidence measures is inconsistent across settings. Through our evaluation framework and analysis, our study provides practical guidance toward developing and evaluating models that selectively use CoT.
title Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
topic Computation and Language
url https://arxiv.org/abs/2510.21007