Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
Fuente:
arXiv
Saved in:
| Main Authors: | Vashurin, Roman, Fadeeva, Ekaterina, Vazhentsev, Artem, Rvanova, Lyudmila, Tsvigun, Akim, Vasilev, Daniil, Xing, Rui, Sadallah, Abdelrahman Boda, Grishchenkov, Kirill, Petrakov, Sergey, Panchenko, Alexander, Baldwin, Timothy, Nakov, Preslav, Panov, Maxim, Shelmanov, Artem |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
by: Shelmanov, Artem, et al.
Published: (2025)
by: Shelmanov, Artem, et al.
Published: (2025)
Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
by: Vazhentsev, Artem, et al.
Published: (2024)
by: Vazhentsev, Artem, et al.
Published: (2024)
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
by: Fadeeva, Ekaterina, et al.
Published: (2025)
by: Fadeeva, Ekaterina, et al.
Published: (2025)
Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval Augmented Generation
by: Fadeeva, Ekaterina, et al.
Published: (2025)
by: Fadeeva, Ekaterina, et al.
Published: (2025)
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
by: Fadeeva, Ekaterina, et al.
Published: (2024)
by: Fadeeva, Ekaterina, et al.
Published: (2024)
Uncertainty Quantification for LLMs through Minimum Bayes Risk: Bridging Confidence and Consistency
by: Vashurin, Roman, et al.
Published: (2025)
by: Vashurin, Roman, et al.
Published: (2025)
Uncertainty Quantification for Large Language Diffusion Models
by: Vazhentsev, Artem, et al.
Published: (2026)
by: Vazhentsev, Artem, et al.
Published: (2026)
Uncertainty-aware abstention in medical diagnosis based on medical texts
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
ReDAct: Uncertainty-Aware Deferral for LLM Agents
by: Piatrashyn, Dzianis, et al.
Published: (2026)
by: Piatrashyn, Dzianis, et al.
Published: (2026)
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
by: Vashurin, Roman, et al.
Published: (2025)
by: Vashurin, Roman, et al.
Published: (2025)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
by: Goloburda, Maiya, et al.
Published: (2026)
by: Goloburda, Maiya, et al.
Published: (2026)
ATGen: A Framework for Active Text Generation
by: Tsvigun, Akim, et al.
Published: (2025)
by: Tsvigun, Akim, et al.
Published: (2025)
M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
by: Rodkin, Ivan, et al.
Published: (2025)
by: Rodkin, Ivan, et al.
Published: (2025)
What Makes Cryptic Crosswords Challenging for LLMs?
by: Sadallah, Abdelrahman, et al.
Published: (2024)
by: Sadallah, Abdelrahman, et al.
Published: (2024)
Are LLMs Good Cryptic Crossword Solvers?
by: Sadallah, Abdelrahman, et al.
Published: (2024)
by: Sadallah, Abdelrahman, et al.
Published: (2024)
SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
Inference-Time Selective Debiasing to Enhance Fairness in Text Classification Models
by: Kuzmin, Gleb, et al.
Published: (2024)
by: Kuzmin, Gleb, et al.
Published: (2024)
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)
by: Sadallah, Abdelrahman, et al.
Published: (2026)
TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling
by: Gorishniy, Yury, et al.
Published: (2024)
by: Gorishniy, Yury, et al.
Published: (2024)
Mathematical model of thyroid gland functioning as a follicles system
by: Ekaterina Vladimirovna Fadeeva
Published: (2021)
by: Ekaterina Vladimirovna Fadeeva
Published: (2021)
Efficient estimation of parameters in marginals in semiparametric multivariate models
by: Medovikov, Ivan, et al.
Published: (2024)
by: Medovikov, Ivan, et al.
Published: (2024)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
by: Vazhentsev, Artem, et al.
Published: (2026)
by: Vazhentsev, Artem, et al.
Published: (2026)
SSCHA-based evolutionary crystal structure prediction at finite temperatures with account for quantum nuclear motion
by: Poletaev, Daniil, et al.
Published: (2025)
by: Poletaev, Daniil, et al.
Published: (2025)
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
by: Rykov, Elisei, et al.
Published: (2025)
by: Rykov, Elisei, et al.
Published: (2025)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
by: Xing, Rui, et al.
Published: (2025)
by: Xing, Rui, et al.
Published: (2025)
CAMAR: Continuous Actions Multi-Agent Routing
by: Pshenitsyn, Artem, et al.
Published: (2025)
by: Pshenitsyn, Artem, et al.
Published: (2025)
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
On Finetuning Tabular Foundation Models
by: Rubachev, Ivan, et al.
Published: (2025)
by: Rubachev, Ivan, et al.
Published: (2025)
TabDDPM: Modelling Tabular Data with Diffusion Models
by: Kotelnikov, Akim, et al.
Published: (2022)
by: Kotelnikov, Akim, et al.
Published: (2022)
Conceptualizing and Operationalizing Prompt Literacy for English Language Learners
by: Ekaterina Tour, et al.
Published: (2025)
by: Ekaterina Tour, et al.
Published: (2025)
Polygraphic resolutions for operated algebras
by: Liu, Zuan, et al.
Published: (2025)
by: Liu, Zuan, et al.
Published: (2025)
Similar Items
-
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
by: Vazhentsev, Artem, et al.
Published: (2025) -
Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models
by: Vazhentsev, Artem, et al.
Published: (2025) -
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
by: Shelmanov, Artem, et al.
Published: (2025) -
Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models
by: Vazhentsev, Artem, et al.
Published: (2024) -
Don't Throw Away Your Beams: Improving Consistency-based Uncertainties in LLMs via Beam Search
by: Fadeeva, Ekaterina, et al.
Published: (2025)