Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiong, Juming, Guo, Kevin, Ni, Congning, Yan, Chao, Brown, Katherine, Baidya, Avinash, Gao, Xiang, Malin, Bradley, Yin, Zhijun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917350861701120
author Xiong, Juming
Guo, Kevin
Ni, Congning
Yan, Chao
Brown, Katherine
Baidya, Avinash
Gao, Xiang
Malin, Bradley
Yin, Zhijun
author_facet Xiong, Juming
Guo, Kevin
Ni, Congning
Yan, Chao
Brown, Katherine
Baidya, Avinash
Gao, Xiang
Malin, Bradley
Yin, Zhijun
contents Large language models (LLMs) achieve strong reasoning performance through chain-of-thought (CoT) reasoning, yet often generate unnecessarily long reasoning paths that incur high inference cost. Recent self-consistency-based approaches further improve accuracy but require sampling and aggregating multiple reasoning trajectories, leading to substantial additional computational overhead. This paper introduces a confidence-aware decision framework that analyzes a single completed reasoning trajectory to adaptively select between single-path and multi-path reasoning. The framework is trained using sentence-level numeric and linguistic features extracted from intermediate reasoning states in the MedQA dataset and generalizes effectively to MathQA, MedMCQA, and MMLU without additional fine-tuning. Experimental results show that the proposed method maintains accuracy comparable to multi-path baselines while using up to 80\% fewer tokens. These findings demonstrate that reasoning trajectories contain rich signals for uncertainty estimation, enabling a simple, transferable mechanism to balance accuracy and efficiency in LLM reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08999
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
Xiong, Juming
Guo, Kevin
Ni, Congning
Yan, Chao
Brown, Katherine
Baidya, Avinash
Gao, Xiang
Malin, Bradley
Yin, Zhijun
Computation and Language
Large language models (LLMs) achieve strong reasoning performance through chain-of-thought (CoT) reasoning, yet often generate unnecessarily long reasoning paths that incur high inference cost. Recent self-consistency-based approaches further improve accuracy but require sampling and aggregating multiple reasoning trajectories, leading to substantial additional computational overhead. This paper introduces a confidence-aware decision framework that analyzes a single completed reasoning trajectory to adaptively select between single-path and multi-path reasoning. The framework is trained using sentence-level numeric and linguistic features extracted from intermediate reasoning states in the MedQA dataset and generalizes effectively to MathQA, MedMCQA, and MMLU without additional fine-tuning. Experimental results show that the proposed method maintains accuracy comparable to multi-path baselines while using up to 80\% fewer tokens. These findings demonstrate that reasoning trajectories contain rich signals for uncertainty estimation, enabling a simple, transferable mechanism to balance accuracy and efficiency in LLM reasoning.
title Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
topic Computation and Language
url https://arxiv.org/abs/2603.08999