Measuring Aleatoric and Epistemic Uncertainty in LLMs: Empirical Evaluation on ID and OOD QA Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Kevin, Moktar, Subre Abdoul, Li, Jia, Li, Kangshuo, Chen, Feng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
por: Liu, Yaokun, et al.
Publicado: (2026)
por: Liu, Yaokun, et al.
Publicado: (2026)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
por: Chandu, Khyathi Raghavi, et al.
Publicado: (2024)
por: Chandu, Khyathi Raghavi, et al.
Publicado: (2024)
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
por: Xiong, Miao, et al.
Publicado: (2023)
por: Xiong, Miao, et al.
Publicado: (2023)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
por: Zong, Qing, et al.
Publicado: (2024)
por: Zong, Qing, et al.
Publicado: (2024)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
por: Hou, Yutao, et al.
Publicado: (2024)
por: Hou, Yutao, et al.
Publicado: (2024)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
por: Dineen, Jacob, et al.
Publicado: (2025)
por: Dineen, Jacob, et al.
Publicado: (2025)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
por: Guo, Kevin H., et al.
Publicado: (2026)
por: Guo, Kevin H., et al.
Publicado: (2026)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
por: Bakman, Yavuz, et al.
Publicado: (2025)
por: Bakman, Yavuz, et al.
Publicado: (2025)
Federated In-Context LLM Agent Learning
por: Wu, Panlong, et al.
Publicado: (2024)
por: Wu, Panlong, et al.
Publicado: (2024)
Estimating Epistemic and Aleatoric Uncertainty with a Single Model
por: Chan, Matthew A., et al.
Publicado: (2024)
por: Chan, Matthew A., et al.
Publicado: (2024)
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
por: Le, Chenqian, et al.
Publicado: (2025)
por: Le, Chenqian, et al.
Publicado: (2025)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
por: Hu, Nan, et al.
Publicado: (2024)
por: Hu, Nan, et al.
Publicado: (2024)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
por: Murugadoss, Bhuvanashree, et al.
Publicado: (2024)
por: Murugadoss, Bhuvanashree, et al.
Publicado: (2024)
Rethinking Aleatoric and Epistemic Uncertainty
por: Smith, Freddie Bickford, et al.
Publicado: (2024)
por: Smith, Freddie Bickford, et al.
Publicado: (2024)
Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
por: Jin, Mingyu, et al.
Publicado: (2026)
por: Jin, Mingyu, et al.
Publicado: (2026)
An Empirical Analysis of Uncertainty in Large Language Model Evaluations
por: Xie, Qiujie, et al.
Publicado: (2025)
por: Xie, Qiujie, et al.
Publicado: (2025)
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
por: Chen, Junjie, et al.
Publicado: (2025)
por: Chen, Junjie, et al.
Publicado: (2025)
JUCAL: Jointly Calibrating Aleatoric and Epistemic Uncertainty in Classification Tasks
por: Heiss, Jakob, et al.
Publicado: (2026)
por: Heiss, Jakob, et al.
Publicado: (2026)
Community-Aligned Behavior Under Uncertainty: Evidence of Epistemic Stance Transfer in LLMs
por: Gerard, Patrick, et al.
Publicado: (2025)
por: Gerard, Patrick, et al.
Publicado: (2025)
Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks
por: Tong, Chaodong, et al.
Publicado: (2025)
por: Tong, Chaodong, et al.
Publicado: (2025)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
por: Dang, Yunkai, et al.
Publicado: (2024)
por: Dang, Yunkai, et al.
Publicado: (2024)
MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs
por: Fan, Yongqi, et al.
Publicado: (2025)
por: Fan, Yongqi, et al.
Publicado: (2025)
Rethinking Epistemic and Aleatoric Uncertainty for Active Open-Set Annotation: An Energy-Based Approach
por: Zong, Chen-Chen, et al.
Publicado: (2025)
por: Zong, Chen-Chen, et al.
Publicado: (2025)
Predictive Uncertainty Quantification for Bird's Eye View Segmentation: A Benchmark and Novel Loss Function
por: Yu, Linlin, et al.
Publicado: (2024)
por: Yu, Linlin, et al.
Publicado: (2024)
Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
por: Zhang, Kaiwei, et al.
Publicado: (2025)
por: Zhang, Kaiwei, et al.
Publicado: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
por: Pan, Wenbo, et al.
Publicado: (2025)
por: Pan, Wenbo, et al.
Publicado: (2025)
Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs
por: Gupta, Ayush, et al.
Publicado: (2025)
por: Gupta, Ayush, et al.
Publicado: (2025)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
por: Lee, Dongryeol, et al.
Publicado: (2024)
por: Lee, Dongryeol, et al.
Publicado: (2024)
BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs via Prompting in Low-Resource QA
por: Jonker, Richard A. A., et al.
Publicado: (2026)
por: Jonker, Richard A. A., et al.
Publicado: (2026)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
por: Bavaresco, Anna, et al.
Publicado: (2024)
por: Bavaresco, Anna, et al.
Publicado: (2024)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
por: Datta, Joyeeta, et al.
Publicado: (2025)
por: Datta, Joyeeta, et al.
Publicado: (2025)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
por: Lai, Viet Dac, et al.
Publicado: (2024)
por: Lai, Viet Dac, et al.
Publicado: (2024)
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
por: Zhang, Jiajie, et al.
Publicado: (2024)
por: Zhang, Jiajie, et al.
Publicado: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
por: Wei, Jianhui, et al.
Publicado: (2025)
por: Wei, Jianhui, et al.
Publicado: (2025)
SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
por: Wu, Mengsong, et al.
Publicado: (2025)
por: Wu, Mengsong, et al.
Publicado: (2025)
The Battle of LLMs: A Comparative Study in Conversational QA Tasks
por: Rangapur, Aryan, et al.
Publicado: (2024)
por: Rangapur, Aryan, et al.
Publicado: (2024)
SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling
por: Sun, Quanen, et al.
Publicado: (2026)
por: Sun, Quanen, et al.
Publicado: (2026)
What External Knowledge is Preferred by LLMs? Characterizing and Exploring Chain of Evidence in Imperfect Context for Multi-Hop QA
por: Chang, Zhiyuan, et al.
Publicado: (2024)
por: Chang, Zhiyuan, et al.
Publicado: (2024)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
por: Li, Kaiyuan, et al.
Publicado: (2026)
por: Li, Kaiyuan, et al.
Publicado: (2026)
JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
por: Saadi, Hossain Shaikh, et al.
Publicado: (2025)
por: Saadi, Hossain Shaikh, et al.
Publicado: (2025)
Ejemplares similares
-
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
por: Liu, Yaokun, et al.
Publicado: (2026) -
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
por: Chandu, Khyathi Raghavi, et al.
Publicado: (2024) -
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
por: Xiong, Miao, et al.
Publicado: (2023) -
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
por: Zong, Qing, et al.
Publicado: (2024) -
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
por: Hou, Yutao, et al.
Publicado: (2024)