NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiayu, Wang, Rui, Zong, Qing, Wang, Yumeng, Qian, Cheng, Zeng, Qingcheng, Zheng, Tianshi, Shi, Haochen, Guo, Dadi, Xu, Baixuan, Li, Chunyang, Song, Yangqiu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025)
by: Zong, Qing, et al.
Published: (2025)
Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
by: Shi, Haochen, et al.
Published: (2025)
by: Shi, Haochen, et al.
Published: (2025)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024)
by: Zong, Qing, et al.
Published: (2024)
Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
by: Liu, Jiayu, et al.
Published: (2025)
by: Liu, Jiayu, et al.
Published: (2025)
Patterns Over Principles: The Fragility of Inductive Reasoning in LLMs under Noisy Observations
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
by: Mo, Yunxiang, et al.
Published: (2025)
by: Mo, Yunxiang, et al.
Published: (2025)
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
by: Xuan, Weihao, et al.
Published: (2025)
by: Xuan, Weihao, et al.
Published: (2025)
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
by: Zheng, Tianshi, et al.
Published: (2024)
by: Zheng, Tianshi, et al.
Published: (2024)
Calibrating Verbalized Confidence with Self-Generated Distractors
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
by: Obadinma, Stephen, et al.
Published: (2025)
by: Obadinma, Stephen, et al.
Published: (2025)
Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning
by: Fan, Wei, et al.
Published: (2026)
by: Fan, Wei, et al.
Published: (2026)
Legal Rule Induction: Towards Generalizable Principle Discovery from Analogous Judicial Precedents
by: Fan, Wei, et al.
Published: (2025)
by: Fan, Wei, et al.
Published: (2025)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
by: Xu, Baixuan, et al.
Published: (2025)
by: Xu, Baixuan, et al.
Published: (2025)
On Verbalized Confidence Scores for LLMs
by: Yang, Daniel, et al.
Published: (2024)
by: Yang, Daniel, et al.
Published: (2024)
EcomEdit: An Automated E-commerce Knowledge Editing Framework for Enhanced Product and Purchase Intention Understanding
by: Lau, Ching Ming Samuel, et al.
Published: (2024)
by: Lau, Ching Ming Samuel, et al.
Published: (2024)
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Double-Calibration: Towards Reliable LLMs via Calibrating Knowledge and Reasoning Confidence
by: Lu, Yuyin, et al.
Published: (2026)
by: Lu, Yuyin, et al.
Published: (2026)
NAACL2025 Tutorial: Adaptation of Large Language Models
by: Ke, Zixuan, et al.
Published: (2025)
by: Ke, Zixuan, et al.
Published: (2025)
Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
by: Zhao, Yunpu, et al.
Published: (2025)
by: Zhao, Yunpu, et al.
Published: (2025)
CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense Reasoning
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
by: Zheng, Zihao, et al.
Published: (2024)
by: Zheng, Zihao, et al.
Published: (2024)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
by: Tsang, Hong Ting, et al.
Published: (2025)
by: Tsang, Hong Ting, et al.
Published: (2025)
Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
by: Wang, Jiecong, et al.
Published: (2026)
by: Wang, Jiecong, et al.
Published: (2026)
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
by: Xuan, Weihao, et al.
Published: (2026)
by: Xuan, Weihao, et al.
Published: (2026)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
by: Wu, Daiqing, et al.
Published: (2025)
by: Wu, Daiqing, et al.
Published: (2025)
Wired for Overconfidence: A Mechanistic Perspective on Inflated Verbalized Confidence in LLMs
by: Zhao, Tianyi, et al.
Published: (2026)
by: Zhao, Tianyi, et al.
Published: (2026)
Fact or Facsimile? Evaluating the Factual Robustness of Modern Retrievers
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
MIKO: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense Discovery
by: Lu, Feihong, et al.
Published: (2024)
by: Lu, Feihong, et al.
Published: (2024)
IDEA: An Interpretable and Editable Decision-Making Framework for LLMs via Verbal-to-Numeric Calibration
by: He, Yanji, et al.
Published: (2026)
by: He, Yanji, et al.
Published: (2026)
Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation
by: Zollo, Thomas, et al.
Published: (2026)
by: Zollo, Thomas, et al.
Published: (2026)
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
by: Chen, Sijia, et al.
Published: (2025)
by: Chen, Sijia, et al.
Published: (2025)
Calibrating Verbalized Probabilities for Large Language Models
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
by: Zhang, Glenn, et al.
Published: (2025)
by: Zhang, Glenn, et al.
Published: (2025)
Confidence-Aware Multi-Field Model Calibration
by: Zhao, Yuang, et al.
Published: (2024)
by: Zhao, Yuang, et al.
Published: (2024)
Similar Items
-
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
by: Zong, Qing, et al.
Published: (2025) -
Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty
by: Wang, Rui, et al.
Published: (2025) -
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
by: Shi, Haochen, et al.
Published: (2025) -
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
by: Zong, Qing, et al.
Published: (2024) -
Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
by: Liu, Jiayu, et al.
Published: (2025)