Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Tianyi, Medina, Johanne, Chawla, Sanjay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Confabulation: The Surprising Value of Large Language Model Hallucinations
by: Sui, Peiqi, et al.
Published: (2024)
by: Sui, Peiqi, et al.
Published: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
by: Smith, Matthew L., et al.
Published: (2026)
by: Smith, Matthew L., et al.
Published: (2026)
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
by: Srey, Ponhvoan, et al.
Published: (2026)
by: Srey, Ponhvoan, et al.
Published: (2026)
Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation
by: Lin, Qinhong, et al.
Published: (2024)
by: Lin, Qinhong, et al.
Published: (2024)
Estimating LLM Uncertainty with Evidence
by: Ma, Huan, et al.
Published: (2025)
by: Ma, Huan, et al.
Published: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Uncertainty Estimation of Large Language Models in Medical Question Answering
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Estimating the Error of Large Language Models at Pairwise Text Comparison
by: Li, Tianyi
Published: (2025)
by: Li, Tianyi
Published: (2025)
The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
by: Ratnakar, Shivam, et al.
Published: (2025)
by: Ratnakar, Shivam, et al.
Published: (2025)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
by: Chatrath, Veronica, et al.
Published: (2024)
by: Chatrath, Veronica, et al.
Published: (2024)
Uncertainty-Aware Dynamic Knowledge Graphs for Reliable Question Answering
by: Takahashi, Yu, et al.
Published: (2025)
by: Takahashi, Yu, et al.
Published: (2025)
T-RAG: Lessons from the LLM Trenches
by: Fatehkia, Masoomali, et al.
Published: (2024)
by: Fatehkia, Masoomali, et al.
Published: (2024)
Can Stories Help LLMs Reason? Curating Information Space Through Narrative
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
by: Javadi, Vahid Sadiri, et al.
Published: (2024)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Can Large Language Models Express Uncertainty Like Human?
by: Tao, Linwei, et al.
Published: (2025)
by: Tao, Linwei, et al.
Published: (2025)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
by: Bajpai, Prasoon, et al.
Published: (2024)
by: Bajpai, Prasoon, et al.
Published: (2024)
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
by: Srivatsa, KV Aditya, et al.
Published: (2025)
by: Srivatsa, KV Aditya, et al.
Published: (2025)
Robust Search with Uncertainty-Aware Value Models for Language Model Reasoning
by: Yu, Fei, et al.
Published: (2025)
by: Yu, Fei, et al.
Published: (2025)
PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
Persona Inconstancy in Multi-Agent LLM Collaboration: Conformity, Confabulation, and Impersonation
by: Baltaji, Razan, et al.
Published: (2024)
by: Baltaji, Razan, et al.
Published: (2024)
Token-Level Uncertainty-Aware Objective for Language Model Post-Training
by: Liu, Tingkai, et al.
Published: (2025)
by: Liu, Tingkai, et al.
Published: (2025)
Breaking Language Barriers: Equitable Performance in Multilingual Language Models
by: Nagar, Tanay, et al.
Published: (2025)
by: Nagar, Tanay, et al.
Published: (2025)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
by: Wu, Tianyi, et al.
Published: (2025)
by: Wu, Tianyi, et al.
Published: (2025)
Can AI-Generated Text be Reliably Detected?
by: Sadasivan, Vinu Sankar, et al.
Published: (2023)
by: Sadasivan, Vinu Sankar, et al.
Published: (2023)
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
An Evaluation of Estimative Uncertainty in Large Language Models
by: Tang, Zhisheng, et al.
Published: (2024)
by: Tang, Zhisheng, et al.
Published: (2024)
Accounting for Sycophancy in Language Model Uncertainty Estimation
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Critical Confabulation: Can LLMs Hallucinate for Social Good?
by: Sui, Peiqi, et al.
Published: (2025)
by: Sui, Peiqi, et al.
Published: (2025)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
by: Wang, Yinong Oliver, et al.
Published: (2025)
by: Wang, Yinong Oliver, et al.
Published: (2025)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
by: Gogoulou, Evangelia, et al.
Published: (2025)
by: Gogoulou, Evangelia, et al.
Published: (2025)
Cleanse: Uncertainty Estimation Approach Using Clustering-based Semantic Consistency in LLMs
by: Joo, Minsuh, et al.
Published: (2025)
by: Joo, Minsuh, et al.
Published: (2025)
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
by: Chen, Sijia, et al.
Published: (2026)
by: Chen, Sijia, et al.
Published: (2026)
ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
by: Zhang, Zhenliang, et al.
Published: (2025)
by: Zhang, Zhenliang, et al.
Published: (2025)
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
by: Chopra, Muskaan, et al.
Published: (2026)
by: Chopra, Muskaan, et al.
Published: (2026)
Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions
by: Baidya, Madhav S., et al.
Published: (2026)
by: Baidya, Madhav S., et al.
Published: (2026)
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026)
by: Myrzakhan, Aidar, et al.
Published: (2026)
Can LLMs Convert Graphs to Text-Attributed Graphs?
by: Wang, Zehong, et al.
Published: (2024)
by: Wang, Zehong, et al.
Published: (2024)
Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models
by: Kauf, Carina, et al.
Published: (2024)
by: Kauf, Carina, et al.
Published: (2024)
Estimation of Concept Explanations Should be Uncertainty Aware
by: Piratla, Vihari, et al.
Published: (2023)
by: Piratla, Vihari, et al.
Published: (2023)
Similar Items
-
Confabulation: The Surprising Value of Large Language Model Hallucinations
by: Sui, Peiqi, et al.
Published: (2024) -
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
by: Smith, Matthew L., et al.
Published: (2026) -
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
by: Srey, Ponhvoan, et al.
Published: (2026) -
Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation
by: Lin, Qinhong, et al.
Published: (2024) -
Estimating LLM Uncertainty with Evidence
by: Ma, Huan, et al.
Published: (2025)