Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Tianyi, He, Yinhan, Zheng, Wendy, Zhang, Yujie, Chen, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mechanistic Circuit-Based Knowledge Editing in Large Language Models
von: Zhao, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhao, Tianyi, et al.
Veröffentlicht: (2026)
On Verbalized Confidence Scores for LLMs
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
von: Obadinma, Stephen, et al.
Veröffentlicht: (2025)
von: Obadinma, Stephen, et al.
Veröffentlicht: (2025)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
von: Li, Chen, et al.
Veröffentlicht: (2026)
von: Li, Chen, et al.
Veröffentlicht: (2026)
How do LLMs Compute Verbal Confidence
von: Kumaran, Dharshan, et al.
Veröffentlicht: (2026)
von: Kumaran, Dharshan, et al.
Veröffentlicht: (2026)
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
von: Zhang, Glenn, et al.
Veröffentlicht: (2025)
von: Zhang, Glenn, et al.
Veröffentlicht: (2025)
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models
von: Groot, Tobias, et al.
Veröffentlicht: (2024)
von: Groot, Tobias, et al.
Veröffentlicht: (2024)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
von: Leng, Jixuan, et al.
Veröffentlicht: (2024)
von: Leng, Jixuan, et al.
Veröffentlicht: (2024)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
von: Yang, Haoyan, et al.
Veröffentlicht: (2024)
von: Yang, Haoyan, et al.
Veröffentlicht: (2024)
ADVICE: Answer-Dependent Verbalized Confidence Estimation
von: Seo, Ki Jung, et al.
Veröffentlicht: (2025)
von: Seo, Ki Jung, et al.
Veröffentlicht: (2025)
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
von: Chhikara, Prateek
Veröffentlicht: (2025)
von: Chhikara, Prateek
Veröffentlicht: (2025)
Are LLM Decisions Faithful to Verbal Confidence?
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
von: Wang, Jiawei, et al.
Veröffentlicht: (2026)
Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
von: Chen, Shiting, et al.
Veröffentlicht: (2025)
von: Chen, Shiting, et al.
Veröffentlicht: (2025)
Calibrating Verbalized Confidence with Self-Generated Distractors
von: Wang, Victor, et al.
Veröffentlicht: (2025)
von: Wang, Victor, et al.
Veröffentlicht: (2025)
Explaining Graph Neural Networks with Large Language Models: A Counterfactual Perspective for Molecular Property Prediction
von: He, Yinhan, et al.
Veröffentlicht: (2024)
von: He, Yinhan, et al.
Veröffentlicht: (2024)
Are Large Language Models More Honest in Their Probabilistic or Verbalized Confidence?
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
von: Ni, Shiyu, et al.
Veröffentlicht: (2024)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
von: He, Yinhan, et al.
Veröffentlicht: (2025)
von: He, Yinhan, et al.
Veröffentlicht: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge
von: Kale, Sahil
Veröffentlicht: (2025)
von: Kale, Sahil
Veröffentlicht: (2025)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
Reforming the Mechanism: Editing Reasoning Patterns in LLMs with Circuit Reshaping
von: Lei, Zhenyu, et al.
Veröffentlicht: (2026)
von: Lei, Zhenyu, et al.
Veröffentlicht: (2026)
IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning
von: He, Yinhan, et al.
Veröffentlicht: (2026)
von: He, Yinhan, et al.
Veröffentlicht: (2026)
Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models
von: Men, Tianyi, et al.
Veröffentlicht: (2024)
von: Men, Tianyi, et al.
Veröffentlicht: (2024)
Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
von: Guan, Jiannan, et al.
Veröffentlicht: (2025)
von: Guan, Jiannan, et al.
Veröffentlicht: (2025)
Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
von: Wang, Ante, et al.
Veröffentlicht: (2025)
von: Wang, Ante, et al.
Veröffentlicht: (2025)
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
von: Li, Yibo, et al.
Veröffentlicht: (2025)
von: Li, Yibo, et al.
Veröffentlicht: (2025)
LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
von: Zhang, Caiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Caiqi, et al.
Veröffentlicht: (2025)
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
von: Niu, Jingcheng, et al.
Veröffentlicht: (2025)
von: Niu, Jingcheng, et al.
Veröffentlicht: (2025)
Verbal Confidence Saturation in 3-9B Open-Weight Instruction-Tuned LLMs: A Pre-Registered Psychometric Validity Screen
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
von: Cacioli, Jon-Paul
Veröffentlicht: (2026)
Supervised Optimism Correction: Be Confident When LLMs Are Sure
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
von: Lu, Junyu, et al.
Veröffentlicht: (2025)
An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages
von: Lu, Yinhan, et al.
Veröffentlicht: (2026)
von: Lu, Yinhan, et al.
Veröffentlicht: (2026)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
von: Xiong, Miao, et al.
Veröffentlicht: (2023)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
von: Zhang, Bonan, et al.
Veröffentlicht: (2025)
von: Zhang, Bonan, et al.
Veröffentlicht: (2025)
Verbalizing LLMs' assumptions to explain and control sycophancy
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation
von: Man, Zhibo, et al.
Veröffentlicht: (2025)
von: Man, Zhibo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mechanistic Circuit-Based Knowledge Editing in Large Language Models
von: Zhao, Tianyi, et al.
Veröffentlicht: (2026) -
On Verbalized Confidence Scores for LLMs
von: Yang, Daniel, et al.
Veröffentlicht: (2024) -
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
von: Obadinma, Stephen, et al.
Veröffentlicht: (2025) -
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
von: Xia, Yuxi, et al.
Veröffentlicht: (2026) -
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
von: Liu, Jiayu, et al.
Veröffentlicht: (2026)