Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
Fuente:
arXiv
Saved in:
| Main Authors: | Lewis-Lim, Samuel, Tan, Xingwei, Zhao, Zhixue, Aletras, Nikolaos |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
by: Lewis-Lim, Samuel, et al.
Published: (2025)
by: Lewis-Lim, Samuel, et al.
Published: (2025)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023)
by: Zhao, Zhixue, et al.
Published: (2023)
Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
by: Alajrami, Ahmed, et al.
Published: (2025)
by: Alajrami, Ahmed, et al.
Published: (2025)
Where does output diversity collapse in post-training?
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
by: Chrysostomou, George, et al.
Published: (2023)
by: Chrysostomou, George, et al.
Published: (2023)
Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision
by: Tan, Xingwei, et al.
Published: (2025)
by: Tan, Xingwei, et al.
Published: (2025)
How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
by: Yamaguchi, Atsuki, et al.
Published: (2024)
by: Yamaguchi, Atsuki, et al.
Published: (2024)
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
by: Villegas, Danae Sánchez, et al.
Published: (2026)
by: Villegas, Danae Sánchez, et al.
Published: (2026)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
by: Emmons, Scott, et al.
Published: (2025)
by: Emmons, Scott, et al.
Published: (2025)
Vocabulary-level Memory Efficiency for Language Model Fine-tuning
by: Williams, Miles, et al.
Published: (2023)
by: Williams, Miles, et al.
Published: (2023)
On the Impact of Calibration Data in Post-training Quantization and Pruning
by: Williams, Miles, et al.
Published: (2023)
by: Williams, Miles, et al.
Published: (2023)
Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models
by: Tan, Xingwei, et al.
Published: (2026)
by: Tan, Xingwei, et al.
Published: (2026)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
by: Permadi, Vynska Amalia, et al.
Published: (2026)
by: Permadi, Vynska Amalia, et al.
Published: (2026)
Fundamental Reasoning Paradigms Induce Out-of-Domain Generalization in Language Models
by: Cao, Mingzi, et al.
Published: (2026)
by: Cao, Mingzi, et al.
Published: (2026)
Large Language Models Decide Early and Explain Later
by: Datta, Ayan, et al.
Published: (2026)
by: Datta, Ayan, et al.
Published: (2026)
Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
by: Xiong, Juming, et al.
Published: (2026)
by: Xiong, Juming, et al.
Published: (2026)
Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates
by: Yamaguchi, Atsuki, et al.
Published: (2025)
by: Yamaguchi, Atsuki, et al.
Published: (2025)
Progressive Depth Up-scaling via Optimal Transport
by: Cao, Mingzi, et al.
Published: (2025)
by: Cao, Mingzi, et al.
Published: (2025)
Self-calibration for Language Model Quantization and Pruning
by: Williams, Miles, et al.
Published: (2024)
by: Williams, Miles, et al.
Published: (2024)
Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
by: Yamaguchi, Atsuki, et al.
Published: (2026)
by: Yamaguchi, Atsuki, et al.
Published: (2026)
How Private are Language Models in Abstractive Summarization?
by: Hughes, Anthony, et al.
Published: (2024)
by: Hughes, Anthony, et al.
Published: (2024)
An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference
by: Yamaguchi, Atsuki, et al.
Published: (2024)
by: Yamaguchi, Atsuki, et al.
Published: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs
by: Tan, Jiejun, et al.
Published: (2024)
by: Tan, Jiejun, et al.
Published: (2024)
Explanation Generation for Contradiction Reconciliation with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
by: Chan, Jason, et al.
Published: (2025)
by: Chan, Jason, et al.
Published: (2025)
GreekBarBench: A Challenging Benchmark for Free-Text Legal Reasoning and Citations
by: Chlapanis, Odysseas S., et al.
Published: (2025)
by: Chlapanis, Odysseas S., et al.
Published: (2025)
Compressing Language Models for Specialized Domains
by: Williams, Miles, et al.
Published: (2025)
by: Williams, Miles, et al.
Published: (2025)
Examining the Limitations of Computational Rumor Detection Models Trained on Static Datasets
by: Mu, Yida, et al.
Published: (2023)
by: Mu, Yida, et al.
Published: (2023)
Enhancing Data Quality through Simple De-duplication: Navigating Responsible Computational Social Science Research
by: Mu, Yida, et al.
Published: (2024)
by: Mu, Yida, et al.
Published: (2024)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025)
by: Xue, Huiyin, et al.
Published: (2025)
It's All About In-Context Learning! Teaching Extremely Low-Resource Languages to LLMs
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
by: Chan, Jason, et al.
Published: (2024)
by: Chan, Jason, et al.
Published: (2024)
Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs
by: Chan, Jason, et al.
Published: (2026)
by: Chan, Jason, et al.
Published: (2026)
Tracing and Reversing Edits in LLMs
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
Supervised Optimism Correction: Be Confident When LLMs Are Sure
by: Zhang, Junjie, et al.
Published: (2025)
by: Zhang, Junjie, et al.
Published: (2025)
When to Trust LLMs: Aligning Confidence with Response Quality
by: Tao, Shuchang, et al.
Published: (2024)
by: Tao, Shuchang, et al.
Published: (2024)
ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
by: Hu, Minda, et al.
Published: (2026)
by: Hu, Minda, et al.
Published: (2026)
Similar Items
-
Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
by: Lewis-Lim, Samuel, et al.
Published: (2025) -
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024) -
Incorporating Attribution Importance for Improving Faithfulness Metrics
by: Zhao, Zhixue, et al.
Published: (2023) -
Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
by: Alajrami, Ahmed, et al.
Published: (2025) -
Where does output diversity collapse in post-training?
by: Karouzos, Constantinos, et al.
Published: (2026)