Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination
Fuente:
arXiv
Guardado en:
| Autores principales: | Agnimo, Yedidia, Korba, Anna, Blangero, Annabelle, Chesneau, Nicolas, Alahari, Karteek |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generalized Discrete Diffusion from Snapshots
por: Zekri, Oussama, et al.
Publicado: (2026)
por: Zekri, Oussama, et al.
Publicado: (2026)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
por: Berrada, Tariq, et al.
Publicado: (2023)
por: Berrada, Tariq, et al.
Publicado: (2023)
On the Shift Invariance of Max Pooling Feature Maps in Convolutional Neural Networks
por: Leterme, Hubert, et al.
Publicado: (2022)
por: Leterme, Hubert, et al.
Publicado: (2022)
From CNNs to Shift-Invariant Twin Models Based on Complex Wavelets
por: Leterme, Hubert, et al.
Publicado: (2022)
por: Leterme, Hubert, et al.
Publicado: (2022)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
por: Sansford, Hannah, et al.
Publicado: (2024)
por: Sansford, Hannah, et al.
Publicado: (2024)
Steer LLM Latents for Hallucination Detection
por: Park, Seongheon, et al.
Publicado: (2025)
por: Park, Seongheon, et al.
Publicado: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
por: Kulkarni, Atharva, et al.
Publicado: (2025)
por: Kulkarni, Atharva, et al.
Publicado: (2025)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
por: Ouyang, Jialin
Publicado: (2025)
por: Ouyang, Jialin
Publicado: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
por: Yadkori, Yasin Abbasi, et al.
Publicado: (2024)
On Mitigating Code LLM Hallucinations with API Documentation
por: Jain, Nihal, et al.
Publicado: (2024)
por: Jain, Nihal, et al.
Publicado: (2024)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
por: Park, Jin Hyun, et al.
Publicado: (2025)
por: Park, Jin Hyun, et al.
Publicado: (2025)
Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models
por: Dey, Prasenjit, et al.
Publicado: (2025)
por: Dey, Prasenjit, et al.
Publicado: (2025)
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
por: Wei, Jiaheng, et al.
Publicado: (2024)
por: Wei, Jiaheng, et al.
Publicado: (2024)
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
por: Miyai, Atsuyuki, et al.
Publicado: (2026)
por: Miyai, Atsuyuki, et al.
Publicado: (2026)
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
por: Duan, Jinhao, et al.
Publicado: (2023)
por: Duan, Jinhao, et al.
Publicado: (2023)
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
por: Hu, Mengya, et al.
Publicado: (2024)
por: Hu, Mengya, et al.
Publicado: (2024)
Hallucination Detection in LLMs: Fast and Memory-Efficient Fine-Tuned Models
por: Arteaga, Gabriel Y., et al.
Publicado: (2024)
por: Arteaga, Gabriel Y., et al.
Publicado: (2024)
ClimateQ&A: Bridging the gap between climate scientists and the general public
por: De La Calzada, Natalia, et al.
Publicado: (2024)
por: De La Calzada, Natalia, et al.
Publicado: (2024)
Are Hallucinations Bad Estimations?
por: Liu, Hude, et al.
Publicado: (2025)
por: Liu, Hude, et al.
Publicado: (2025)
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
por: Divekar, Abhishek, et al.
Publicado: (2026)
por: Divekar, Abhishek, et al.
Publicado: (2026)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
por: Mündler, Niels, et al.
Publicado: (2023)
por: Mündler, Niels, et al.
Publicado: (2023)
Estimation of Concept Explanations Should be Uncertainty Aware
por: Piratla, Vihari, et al.
Publicado: (2023)
por: Piratla, Vihari, et al.
Publicado: (2023)
The Phenomenology of Hallucinations
por: Ruscio, Valeria, et al.
Publicado: (2026)
por: Ruscio, Valeria, et al.
Publicado: (2026)
How Uncertainty Estimation Scales with Sampling in Reasoning Models
por: Del, Maksym, et al.
Publicado: (2026)
por: Del, Maksym, et al.
Publicado: (2026)
LLM Chemistry Estimation for Multi-LLM Recommendation
por: Sanchez, Huascar, et al.
Publicado: (2025)
por: Sanchez, Huascar, et al.
Publicado: (2025)
Consistency Is the Key: Detecting Hallucinations in LLM Generated Text By Checking Inconsistencies About Key Facts
por: Gupta, Raavi, et al.
Publicado: (2025)
por: Gupta, Raavi, et al.
Publicado: (2025)
Revisiting Uncertainty Estimation and Calibration of Large Language Models
por: Tao, Linwei, et al.
Publicado: (2025)
por: Tao, Linwei, et al.
Publicado: (2025)
Towards Multilingual LLM Evaluation for European Languages
por: Thellmann, Klaudia, et al.
Publicado: (2024)
por: Thellmann, Klaudia, et al.
Publicado: (2024)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
por: Rahman, Subhey Sadi, et al.
Publicado: (2025)
por: Rahman, Subhey Sadi, et al.
Publicado: (2025)
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization
por: Tang, Zilu, et al.
Publicado: (2025)
por: Tang, Zilu, et al.
Publicado: (2025)
Beyond Relevance: Utility-Centric Retrieval in the LLM Era
por: Zhang, Hengran, et al.
Publicado: (2026)
por: Zhang, Hengran, et al.
Publicado: (2026)
Deep Learning-based Prediction of Clinical Trial Enrollment with Uncertainty Estimates
por: Do, Tien Huu, et al.
Publicado: (2025)
por: Do, Tien Huu, et al.
Publicado: (2025)
Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?
por: Hashimoto, Wataru, et al.
Publicado: (2024)
por: Hashimoto, Wataru, et al.
Publicado: (2024)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
por: Guo, Kevin H., et al.
Publicado: (2026)
por: Guo, Kevin H., et al.
Publicado: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
por: Liu, Mingyi
Publicado: (2026)
por: Liu, Mingyi
Publicado: (2026)
GENUINE: Graph Enhanced Multi-level Uncertainty Estimation for Large Language Models
por: Wang, Tuo, et al.
Publicado: (2025)
por: Wang, Tuo, et al.
Publicado: (2025)
Efficient Nearest Neighbor based Uncertainty Estimation for Natural Language Processing Tasks
por: Hashimoto, Wataru, et al.
Publicado: (2024)
por: Hashimoto, Wataru, et al.
Publicado: (2024)
Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation
por: Lee, JoonHo, et al.
Publicado: (2024)
por: Lee, JoonHo, et al.
Publicado: (2024)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
por: Zhang, Tunyu, et al.
Publicado: (2025)
por: Zhang, Tunyu, et al.
Publicado: (2025)
Ejemplares similares
-
Generalized Discrete Diffusion from Snapshots
por: Zekri, Oussama, et al.
Publicado: (2026) -
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
por: Berrada, Tariq, et al.
Publicado: (2023) -
On the Shift Invariance of Max Pooling Feature Maps in Convolutional Neural Networks
por: Leterme, Hubert, et al.
Publicado: (2022) -
From CNNs to Shift-Invariant Twin Models Based on Complex Wavelets
por: Leterme, Hubert, et al.
Publicado: (2022) -
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
por: Sansford, Hannah, et al.
Publicado: (2024)