Uncertainty Quantification for LLMs through Minimum Bayes Risk: Bridging Confidence and Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Vashurin, Roman, Goloburda, Maiya, Ilina, Albina, Rubashevskii, Aleksandr, Nakov, Preslav, Shelmanov, Artem, Panov, Maxim |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
by: Fadeeva, Ekaterina, et al.
Published: (2025)
by: Fadeeva, Ekaterina, et al.
Published: (2025)
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
by: Vashurin, Roman, et al.
Published: (2025)
by: Vashurin, Roman, et al.
Published: (2025)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
by: Goloburda, Maiya, et al.
Published: (2026)
by: Goloburda, Maiya, et al.
Published: (2026)
Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
by: Fadeeva, Ekaterina, et al.
Published: (2025)
by: Fadeeva, Ekaterina, et al.
Published: (2025)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
by: Fadeeva, Ekaterina, et al.
Published: (2024)
by: Fadeeva, Ekaterina, et al.
Published: (2024)
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
by: Vashurin, Roman, et al.
Published: (2024)
by: Vashurin, Roman, et al.
Published: (2024)
Uncertainty Quantification for Large Language Diffusion Models
by: Vazhentsev, Artem, et al.
Published: (2026)
by: Vazhentsev, Artem, et al.
Published: (2026)
Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
by: Vazhentsev, Artem, et al.
Published: (2024)
by: Vazhentsev, Artem, et al.
Published: (2024)
Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
ReDAct: Uncertainty-Aware Deferral for LLM Agents
by: Piatrashyn, Dzianis, et al.
Published: (2026)
by: Piatrashyn, Dzianis, et al.
Published: (2026)
Uncertainty-aware abstention in medical diagnosis based on medical texts
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
by: Laiyk, Nurkhan, et al.
Published: (2025)
by: Laiyk, Nurkhan, et al.
Published: (2025)
Multidimensional Uncertainty Quantification via Optimal Transport
by: Kotelevskii, Nikita, et al.
Published: (2025)
by: Kotelevskii, Nikita, et al.
Published: (2025)
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
by: Shelmanov, Artem, et al.
Published: (2025)
by: Shelmanov, Artem, et al.
Published: (2025)
Enhancing Factuality through Consensus and Consistency in Summarization Using Minimum Bayes Risk Decoding
by: Soetedjo, Riza Setiawan, et al.
Published: (2026)
by: Soetedjo, Riza Setiawan, et al.
Published: (2026)
Vikhr: The Family of Open-Source Instruction-Tuned Large Language Models for Russian
by: Nikolich, Aleksandr, et al.
Published: (2024)
by: Nikolich, Aleksandr, et al.
Published: (2024)
Uncertainty-Aware Decoding with Minimum Bayes Risk
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
A Survey of Confidence Estimation and Calibration in Large Language Models
by: Geng, Jiahui, et al.
Published: (2023)
by: Geng, Jiahui, et al.
Published: (2023)
Qorgau: Evaluating LLM Safety in Kazakh-Russian Bilingual Contexts
by: Goloburda, Maiya, et al.
Published: (2025)
by: Goloburda, Maiya, et al.
Published: (2025)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
by: Togmanov, Mukhammed, et al.
Published: (2025)
by: Togmanov, Mukhammed, et al.
Published: (2025)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
by: Ren, Kaixuan, et al.
Published: (2025)
by: Ren, Kaixuan, et al.
Published: (2025)
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition
by: Alhindi, Tariq, et al.
Published: (2023)
by: Alhindi, Tariq, et al.
Published: (2023)
Rethinking STS and NLI in Large Language Models
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
by: Sundriyal, Megha, et al.
Published: (2023)
by: Sundriyal, Megha, et al.
Published: (2023)
Adapting Fake News Detection to the Era of Large Language Models
by: Su, Jinyan, et al.
Published: (2023)
by: Su, Jinyan, et al.
Published: (2023)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
by: Tomar, Raj Vardhan, et al.
Published: (2025)
by: Tomar, Raj Vardhan, et al.
Published: (2025)
How Does Prefix Matter in Reasoning Model Tuning?
by: Tomar, Raj Vardhan, et al.
Published: (2026)
by: Tomar, Raj Vardhan, et al.
Published: (2026)
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
by: Bates, Luke, et al.
Published: (2025)
by: Bates, Luke, et al.
Published: (2025)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
by: Elozeiri, Kareem, et al.
Published: (2025)
by: Elozeiri, Kareem, et al.
Published: (2025)
Missci: Reconstructing Fallacies in Misrepresented Science
by: Glockner, Max, et al.
Published: (2024)
by: Glockner, Max, et al.
Published: (2024)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
by: Glockner, Max, et al.
Published: (2024)
by: Glockner, Max, et al.
Published: (2024)
Confidence Improves Self-Consistency in LLMs
by: Taubenfeld, Amir, et al.
Published: (2025)
by: Taubenfeld, Amir, et al.
Published: (2025)
Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space
by: Qiu, Xin, et al.
Published: (2024)
by: Qiu, Xin, et al.
Published: (2024)
Structure-Conditional Minimum Bayes Risk Decoding
by: Eikema, Bryan, et al.
Published: (2025)
by: Eikema, Bryan, et al.
Published: (2025)
Similar Items
-
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
by: Fadeeva, Ekaterina, et al.
Published: (2025) -
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
by: Vashurin, Roman, et al.
Published: (2025) -
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
by: Goloburda, Maiya, et al.
Published: (2026) -
Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
by: Fadeeva, Ekaterina, et al.
Published: (2025) -
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
by: Rubashevskii, Aleksandr, et al.
Published: (2026)