Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Wenbo, Padmanabhan, Veena, Giyahchi, Tootiya, Wong, Elaine, Akoglu, Leman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
di: Deshpande, Vijeta, et al.
Pubblicazione: (2026)
di: Deshpande, Vijeta, et al.
Pubblicazione: (2026)
FoMo-0D: A Foundation Model for Zero-shot Tabular Outlier Detection
di: Shen, Yuchen, et al.
Pubblicazione: (2024)
di: Shen, Yuchen, et al.
Pubblicazione: (2024)
Toward Privileged Foundation Models:LUPI for Accelerated and Improved Learning
di: Ding, Xueying, et al.
Pubblicazione: (2026)
di: Ding, Xueying, et al.
Pubblicazione: (2026)
CoBAD: Modeling Collective Behaviors for Human Mobility Anomaly Detection
di: Wen, Haomin, et al.
Pubblicazione: (2025)
di: Wen, Haomin, et al.
Pubblicazione: (2025)
Fast Unsupervised Deep Outlier Model Selection with Hypernetworks
di: Ding, Xueying, et al.
Pubblicazione: (2023)
di: Ding, Xueying, et al.
Pubblicazione: (2023)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
di: Zhou, Tianyang, et al.
Pubblicazione: (2026)
di: Zhou, Tianyang, et al.
Pubblicazione: (2026)
Uncertainty-aware Human Mobility Modeling and Anomaly Detection
di: Wen, Haomin, et al.
Pubblicazione: (2024)
di: Wen, Haomin, et al.
Pubblicazione: (2024)
On the Detection of Reviewer-Author Collusion Rings From Paper Bidding
di: Jecmen, Steven, et al.
Pubblicazione: (2024)
di: Jecmen, Steven, et al.
Pubblicazione: (2024)
Is LLMs Hallucination Usable? LLM-based Negative Reasoning for Fake News Detection
di: Zhang, Chaowei, et al.
Pubblicazione: (2025)
di: Zhang, Chaowei, et al.
Pubblicazione: (2025)
Banishing LLM Hallucinations Requires Rethinking Generalization
di: Li, Johnny, et al.
Pubblicazione: (2024)
di: Li, Johnny, et al.
Pubblicazione: (2024)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
di: Gunjal, Anisha, et al.
Pubblicazione: (2024)
di: Gunjal, Anisha, et al.
Pubblicazione: (2024)
MacrOData: New Benchmarks of Thousands of Datasets for Tabular Outlier Detection
di: Ding, Xueying, et al.
Pubblicazione: (2026)
di: Ding, Xueying, et al.
Pubblicazione: (2026)
AI Hallucinations: A Misnomer Worth Clarifying
di: Maleki, Negar, et al.
Pubblicazione: (2024)
di: Maleki, Negar, et al.
Pubblicazione: (2024)
LettuceDetect: A Hallucination Detection Framework for RAG Applications
di: Kovács, Ádám, et al.
Pubblicazione: (2025)
di: Kovács, Ádám, et al.
Pubblicazione: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
Removal of Hallucination on Hallucination: Debate-Augmented RAG
di: Hu, Wentao, et al.
Pubblicazione: (2025)
di: Hu, Wentao, et al.
Pubblicazione: (2025)
Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
di: Taş, Selva, et al.
Pubblicazione: (2025)
di: Taş, Selva, et al.
Pubblicazione: (2025)
Evaluating Progress in Graph Foundation Models: A Comprehensive Benchmark and New Insights
di: Yu, Xingtong, et al.
Pubblicazione: (2026)
di: Yu, Xingtong, et al.
Pubblicazione: (2026)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
Diagnosing LLM Arbitration Behavior over Pre-evidence Epistemic States in RAG-based Fact-Checking
di: Sun, Yuxi, et al.
Pubblicazione: (2026)
di: Sun, Yuxi, et al.
Pubblicazione: (2026)
Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
di: Du, Xueying, et al.
Pubblicazione: (2024)
di: Du, Xueying, et al.
Pubblicazione: (2024)
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection
di: Xu, Cheng, et al.
Pubblicazione: (2026)
di: Xu, Cheng, et al.
Pubblicazione: (2026)
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025)
di: Bang, Yejin, et al.
Pubblicazione: (2025)
Fine-tuning with RAG for Improving LLM Learning of New Skills
di: Ibrahim, Humaid, et al.
Pubblicazione: (2025)
di: Ibrahim, Humaid, et al.
Pubblicazione: (2025)
What Breaks Knowledge Graph based RAG? Benchmarking and Empirical Insights into Reasoning under Incomplete Knowledge
di: Zhou, Dongzhuoran, et al.
Pubblicazione: (2025)
di: Zhou, Dongzhuoran, et al.
Pubblicazione: (2025)
TPA: Next Token Probability Attribution for Detecting Hallucinations in RAG
di: Lu, Pengqian, et al.
Pubblicazione: (2025)
di: Lu, Pengqian, et al.
Pubblicazione: (2025)
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
di: Luo, Jinyuan, et al.
Pubblicazione: (2025)
di: Luo, Jinyuan, et al.
Pubblicazione: (2025)
Behavior Alignment: A New Perspective of Evaluating LLM-based Conversational Recommender Systems
di: Yang, Dayu, et al.
Pubblicazione: (2024)
di: Yang, Dayu, et al.
Pubblicazione: (2024)
Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching
di: Wei, Rongzhe, et al.
Pubblicazione: (2026)
di: Wei, Rongzhe, et al.
Pubblicazione: (2026)
REFRAG: Rethinking RAG based Decoding
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models
di: Zuo, Kaiwen, et al.
Pubblicazione: (2024)
di: Zuo, Kaiwen, et al.
Pubblicazione: (2024)
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA
di: Baqar, Mohammad, et al.
Pubblicazione: (2025)
di: Baqar, Mohammad, et al.
Pubblicazione: (2025)
OptArgus: A Multi-Agent System to Detect Hallucinations in LLM-based Optimization Modeling
di: Li, Zhong, et al.
Pubblicazione: (2026)
di: Li, Zhong, et al.
Pubblicazione: (2026)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
Design Behaviour Codes (DBCs): A Taxonomy-Driven Layered Governance Benchmark for Large Language Models
di: Mohan, G. Madan, et al.
Pubblicazione: (2026)
di: Mohan, G. Madan, et al.
Pubblicazione: (2026)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
di: Chen, Kedi, et al.
Pubblicazione: (2024)
di: Chen, Kedi, et al.
Pubblicazione: (2024)
HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention
di: Goswami, Saumya, et al.
Pubblicazione: (2025)
di: Goswami, Saumya, et al.
Pubblicazione: (2025)
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
di: Zhao, Feiyu, et al.
Pubblicazione: (2026)
di: Zhao, Feiyu, et al.
Pubblicazione: (2026)
GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation
di: Sorodoc, Ionut-Teodor, et al.
Pubblicazione: (2025)
di: Sorodoc, Ionut-Teodor, et al.
Pubblicazione: (2025)
Hyper-RAG: Combating LLM Hallucinations using Hypergraph-Driven Retrieval-Augmented Generation
di: Feng, Yifan, et al.
Pubblicazione: (2025)
di: Feng, Yifan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
di: Deshpande, Vijeta, et al.
Pubblicazione: (2026) -
FoMo-0D: A Foundation Model for Zero-shot Tabular Outlier Detection
di: Shen, Yuchen, et al.
Pubblicazione: (2024) -
Toward Privileged Foundation Models:LUPI for Accelerated and Improved Learning
di: Ding, Xueying, et al.
Pubblicazione: (2026) -
CoBAD: Modeling Collective Behaviors for Human Mobility Anomaly Detection
di: Wen, Haomin, et al.
Pubblicazione: (2025) -
Fast Unsupervised Deep Outlier Model Selection with Hypernetworks
di: Ding, Xueying, et al.
Pubblicazione: (2023)