Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Agnimo, Yedidia, Korba, Anna, Blangero, Annabelle, Chesneau, Nicolas, Alahari, Karteek |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Generalized Discrete Diffusion from Snapshots
par: Zekri, Oussama, et autres
Publié: (2026)
par: Zekri, Oussama, et autres
Publié: (2026)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
par: Berrada, Tariq, et autres
Publié: (2023)
par: Berrada, Tariq, et autres
Publié: (2023)
On the Shift Invariance of Max Pooling Feature Maps in Convolutional Neural Networks
par: Leterme, Hubert, et autres
Publié: (2022)
par: Leterme, Hubert, et autres
Publié: (2022)
From CNNs to Shift-Invariant Twin Models Based on Complex Wavelets
par: Leterme, Hubert, et autres
Publié: (2022)
par: Leterme, Hubert, et autres
Publié: (2022)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
par: Sansford, Hannah, et autres
Publié: (2024)
par: Sansford, Hannah, et autres
Publié: (2024)
Steer LLM Latents for Hallucination Detection
par: Park, Seongheon, et autres
Publié: (2025)
par: Park, Seongheon, et autres
Publié: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
par: Kulkarni, Atharva, et autres
Publié: (2025)
par: Kulkarni, Atharva, et autres
Publié: (2025)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
par: Ouyang, Jialin
Publié: (2025)
par: Ouyang, Jialin
Publié: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
par: Yadkori, Yasin Abbasi, et autres
Publié: (2024)
par: Yadkori, Yasin Abbasi, et autres
Publié: (2024)
On Mitigating Code LLM Hallucinations with API Documentation
par: Jain, Nihal, et autres
Publié: (2024)
par: Jain, Nihal, et autres
Publié: (2024)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
par: Park, Jin Hyun, et autres
Publié: (2025)
par: Park, Jin Hyun, et autres
Publié: (2025)
Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models
par: Dey, Prasenjit, et autres
Publié: (2025)
par: Dey, Prasenjit, et autres
Publié: (2025)
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
par: Wei, Jiaheng, et autres
Publié: (2024)
par: Wei, Jiaheng, et autres
Publié: (2024)
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
par: Miyai, Atsuyuki, et autres
Publié: (2026)
par: Miyai, Atsuyuki, et autres
Publié: (2026)
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
par: Duan, Jinhao, et autres
Publié: (2023)
par: Duan, Jinhao, et autres
Publié: (2023)
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
par: Hu, Mengya, et autres
Publié: (2024)
par: Hu, Mengya, et autres
Publié: (2024)
Hallucination Detection in LLMs: Fast and Memory-Efficient Fine-Tuned Models
par: Arteaga, Gabriel Y., et autres
Publié: (2024)
par: Arteaga, Gabriel Y., et autres
Publié: (2024)
ClimateQ&A: Bridging the gap between climate scientists and the general public
par: De La Calzada, Natalia, et autres
Publié: (2024)
par: De La Calzada, Natalia, et autres
Publié: (2024)
Are Hallucinations Bad Estimations?
par: Liu, Hude, et autres
Publié: (2025)
par: Liu, Hude, et autres
Publié: (2025)
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
par: Divekar, Abhishek, et autres
Publié: (2026)
par: Divekar, Abhishek, et autres
Publié: (2026)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
par: Mündler, Niels, et autres
Publié: (2023)
par: Mündler, Niels, et autres
Publié: (2023)
Estimation of Concept Explanations Should be Uncertainty Aware
par: Piratla, Vihari, et autres
Publié: (2023)
par: Piratla, Vihari, et autres
Publié: (2023)
The Phenomenology of Hallucinations
par: Ruscio, Valeria, et autres
Publié: (2026)
par: Ruscio, Valeria, et autres
Publié: (2026)
How Uncertainty Estimation Scales with Sampling in Reasoning Models
par: Del, Maksym, et autres
Publié: (2026)
par: Del, Maksym, et autres
Publié: (2026)
LLM Chemistry Estimation for Multi-LLM Recommendation
par: Sanchez, Huascar, et autres
Publié: (2025)
par: Sanchez, Huascar, et autres
Publié: (2025)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
par: Tao, Linwei, et autres
Publié: (2025)
par: Tao, Linwei, et autres
Publié: (2025)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
par: Gupta, Raavi, et autres
Publié: (2025)
par: Gupta, Raavi, et autres
Publié: (2025)
Towards Multilingual LLM Evaluation for European Languages
par: Thellmann, Klaudia, et autres
Publié: (2024)
par: Thellmann, Klaudia, et autres
Publié: (2024)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
par: Rahman, Subhey Sadi, et autres
Publié: (2025)
par: Rahman, Subhey Sadi, et autres
Publié: (2025)
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization
par: Tang, Zilu, et autres
Publié: (2025)
par: Tang, Zilu, et autres
Publié: (2025)
Beyond Relevance: Utility-Centric Retrieval in the LLM Era
par: Zhang, Hengran, et autres
Publié: (2026)
par: Zhang, Hengran, et autres
Publié: (2026)
Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
par: Do, Tien Huu, et autres
Publié: (2025)
par: Do, Tien Huu, et autres
Publié: (2025)
Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?
par: Hashimoto, Wataru, et autres
Publié: (2024)
par: Hashimoto, Wataru, et autres
Publié: (2024)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
par: Guo, Kevin H., et autres
Publié: (2026)
par: Guo, Kevin H., et autres
Publié: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
par: Liu, Yixin, et autres
Publié: (2025)
par: Liu, Yixin, et autres
Publié: (2025)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
par: Liu, Mingyi
Publié: (2026)
par: Liu, Mingyi
Publié: (2026)
GENUINE: Graph Enhanced Multi-level Uncertainty Estimation for Large Language Models
par: Wang, Tuo, et autres
Publié: (2025)
par: Wang, Tuo, et autres
Publié: (2025)
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
par: Hashimoto, Wataru, et autres
Publié: (2024)
par: Hashimoto, Wataru, et autres
Publié: (2024)
Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation
par: Lee, JoonHo, et autres
Publié: (2024)
par: Lee, JoonHo, et autres
Publié: (2024)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
par: Zhang, Tunyu, et autres
Publié: (2025)
par: Zhang, Tunyu, et autres
Publié: (2025)
Documents similaires
-
Generalized Discrete Diffusion from Snapshots
par: Zekri, Oussama, et autres
Publié: (2026) -
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
par: Berrada, Tariq, et autres
Publié: (2023) -
On the Shift Invariance of Max Pooling Feature Maps in Convolutional Neural Networks
par: Leterme, Hubert, et autres
Publié: (2022) -
From CNNs to Shift-Invariant Twin Models Based on Complex Wavelets
par: Leterme, Hubert, et autres
Publié: (2022) -
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
par: Sansford, Hannah, et autres
Publié: (2024)