Are Large Language Models More Honest in Their Probabilistic or Verbalized Confidence?
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, Shiyu, Bi, Keping, Yu, Lulu, Guo, Jiafeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
by: Ni, Shiyu, et al.
Published: (2024)
by: Ni, Shiyu, et al.
Published: (2024)
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)
by: Tang, Minghao, et al.
Published: (2025)
How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality
by: Tu, Minzhu, et al.
Published: (2026)
by: Tu, Minzhu, et al.
Published: (2026)
Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs
by: Ding, Zhikai, et al.
Published: (2025)
by: Ding, Zhikai, et al.
Published: (2025)
Contextual Dual Learning Algorithm with Listwise Distillation for Unbiased Learning to Rank
by: Yu, Lulu, et al.
Published: (2024)
by: Yu, Lulu, et al.
Published: (2024)
Evaluating Implicit Bias in Large Language Models by Attacking From a Psychometric Perspective
by: Wen, Yuchen, et al.
Published: (2024)
by: Wen, Yuchen, et al.
Published: (2024)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Annotation-Efficient Universal Honesty Alignment
by: Ni, Shiyu, et al.
Published: (2025)
by: Ni, Shiyu, et al.
Published: (2025)
MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning
by: Cui, Wanqing, et al.
Published: (2024)
by: Cui, Wanqing, et al.
Published: (2024)
A Comparative Study of Specialized LLMs as Dense Retrievers
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
by: Zhang, Glenn, et al.
Published: (2025)
by: Zhang, Glenn, et al.
Published: (2025)
HonestLLM: Toward an Honest and Helpful Large Language Model
by: Gao, Chujie, et al.
Published: (2024)
by: Gao, Chujie, et al.
Published: (2024)
LINKAGE: Listwise Ranking among Varied-Quality References for Non-Factoid QA Evaluation via LLMs
by: Yang, Sihui, et al.
Published: (2024)
by: Yang, Sihui, et al.
Published: (2024)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
by: Zhang, Hengran, et al.
Published: (2024)
by: Zhang, Hengran, et al.
Published: (2024)
Beyond Relevance: Utility-Centric Retrieval in the LLM Era
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)
by: Tang, Minghao, et al.
Published: (2025)
Estimating Commonsense Plausibility through Semantic Shifts
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
ConfTuner: Training Large Language Models to Express Their Confidence Verbally
by: Li, Yibo, et al.
Published: (2025)
by: Li, Yibo, et al.
Published: (2025)
ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models
by: Li, Chen, et al.
Published: (2026)
by: Li, Chen, et al.
Published: (2026)
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
On Verbalized Confidence Scores for LLMs
by: Yang, Daniel, et al.
Published: (2024)
by: Yang, Daniel, et al.
Published: (2024)
Toward Honest Language Models for Deductive Reasoning
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
Bagging-Based Model Merging for Robust General Text Embeddings
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
ADVICE: Answer-Dependent Verbalized Confidence Estimation
by: Seo, Ki Jung, et al.
Published: (2025)
by: Seo, Ki Jung, et al.
Published: (2025)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
by: Obadinma, Stephen, et al.
Published: (2025)
by: Obadinma, Stephen, et al.
Published: (2025)
Searching for Best Practices in Medical Transcription with Large Language Model
by: Li, Jiafeng, et al.
Published: (2024)
by: Li, Jiafeng, et al.
Published: (2024)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
CIR at the NTCIR-17 ULTRE-2 Task
by: Yu, Lulu, et al.
Published: (2023)
by: Yu, Lulu, et al.
Published: (2023)
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Prism-$Δ$: Differential Subspace Steering for Prompt Highlighting in Large Language Models
by: Ge, Yuyao, et al.
Published: (2026)
by: Ge, Yuyao, et al.
Published: (2026)
Calibrating Verbalized Probabilities for Large Language Models
by: Wang, Cheng, et al.
Published: (2024)
by: Wang, Cheng, et al.
Published: (2024)
Calibrating Verbalized Confidence with Self-Generated Distractors
by: Wang, Victor, et al.
Published: (2025)
by: Wang, Victor, et al.
Published: (2025)
Calibrating the Confidence of Large Language Models by Eliciting Fidelity
by: Zhang, Mozhi, et al.
Published: (2024)
by: Zhang, Mozhi, et al.
Published: (2024)
Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration
by: Wu, Guangxin, et al.
Published: (2026)
by: Wu, Guangxin, et al.
Published: (2026)
NAACL: Noise-AwAre Verbal Confidence Calibration for Robust LLMs in RAG Systems
by: Liu, Jiayu, et al.
Published: (2026)
by: Liu, Jiayu, et al.
Published: (2026)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
Similar Items
-
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025) -
When Do LLMs Need Retrieval Augmentation? Mitigating LLMs' Overconfidence Helps Retrieval Augmentation
by: Ni, Shiyu, et al.
Published: (2024) -
How Knowledge Popularity Influences and Enhances LLM Knowledge Boundary Perception
by: Ni, Shiyu, et al.
Published: (2025) -
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025) -
How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality
by: Tu, Minzhu, et al.
Published: (2026)