KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Robertson, Alex, Liang, Huizhi, Gani, Mahbub, Kumar, Rohit, Rajamohan, Srijith |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
par: Wang, Sizhe, et autres
Publié: (2024)
par: Wang, Sizhe, et autres
Publié: (2024)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
par: Sansford, Hannah, et autres
Publié: (2024)
par: Sansford, Hannah, et autres
Publié: (2024)
D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs
par: Ding, Yue, et autres
Publié: (2025)
par: Ding, Yue, et autres
Publié: (2025)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
par: Markowitz, Elan, et autres
Publié: (2025)
par: Markowitz, Elan, et autres
Publié: (2025)
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
par: Lavrinovics, Ernests, et autres
Publié: (2025)
par: Lavrinovics, Ernests, et autres
Publié: (2025)
Training for Compositional Sensitivity Reduces Dense Retrieval Generalization
par: Ralev, Radoslav, et autres
Publié: (2026)
par: Ralev, Radoslav, et autres
Publié: (2026)
MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models
par: Liu, Weixin, et autres
Publié: (2026)
par: Liu, Weixin, et autres
Publié: (2026)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
par: Kumar, Mahesh, et autres
Publié: (2026)
par: Kumar, Mahesh, et autres
Publié: (2026)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
par: Liu, Zhiqiang, et autres
Publié: (2025)
par: Liu, Zhiqiang, et autres
Publié: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
par: Perlitz, Yotam, et autres
Publié: (2024)
par: Perlitz, Yotam, et autres
Publié: (2024)
From Hallucinations to Facts: Enhancing Language Models with Curated Knowledge Graphs
par: Joshi, Ratnesh Kumar, et autres
Publié: (2024)
par: Joshi, Ratnesh Kumar, et autres
Publié: (2024)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
par: Ghanem, Hussam, et autres
Publié: (2025)
par: Ghanem, Hussam, et autres
Publié: (2025)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
par: Nguyen, Hieu, et autres
Publié: (2025)
par: Nguyen, Hieu, et autres
Publié: (2025)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
par: Chen, Yongrui, et autres
Publié: (2025)
par: Chen, Yongrui, et autres
Publié: (2025)
CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders
par: Gulko, Alex, et autres
Publié: (2025)
par: Gulko, Alex, et autres
Publié: (2025)
Combining LLMs and Knowledge Graphs to Reduce Hallucinations in Question Answering
par: Pusch, Larissa, et autres
Publié: (2024)
par: Pusch, Larissa, et autres
Publié: (2024)
LastingBench: Defend Benchmarks Against Knowledge Leakage
par: Fang, Yixiong, et autres
Publié: (2025)
par: Fang, Yixiong, et autres
Publié: (2025)
Characterizing Knowledge Graph Tasks in LLM Benchmarks Using Cognitive Complexity Frameworks
par: Todorovikj, Sara, et autres
Publié: (2025)
par: Todorovikj, Sara, et autres
Publié: (2025)
Agent-as-a-Graph: Knowledge Graph-Based Tool and Agent Retrieval for LLM Multi-Agent Systems
par: Nizar, Faheem, et autres
Publié: (2025)
par: Nizar, Faheem, et autres
Publié: (2025)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
par: Zhang, Yuji, et autres
Publié: (2025)
par: Zhang, Yuji, et autres
Publié: (2025)
LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoning
par: Xing, Wang, et autres
Publié: (2026)
par: Xing, Wang, et autres
Publié: (2026)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
par: Feng, Shangbin, et autres
Publié: (2024)
par: Feng, Shangbin, et autres
Publié: (2024)
EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
par: Zaporojets, Klim, et autres
Publié: (2025)
par: Zaporojets, Klim, et autres
Publié: (2025)
MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge
par: Fu, Baochen, et autres
Publié: (2026)
par: Fu, Baochen, et autres
Publié: (2026)
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
par: Du, Yuntao, et autres
Publié: (2025)
par: Du, Yuntao, et autres
Publié: (2025)
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
par: Xie, Pengzhen, et autres
Publié: (2025)
par: Xie, Pengzhen, et autres
Publié: (2025)
On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
par: Kalavasis, Alkis, et autres
Publié: (2024)
par: Kalavasis, Alkis, et autres
Publié: (2024)
Grounding LLM Reasoning with Knowledge Graphs
par: Amayuelas, Alfonso, et autres
Publié: (2025)
par: Amayuelas, Alfonso, et autres
Publié: (2025)
Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
par: Agrawal, Garima, et autres
Publié: (2023)
par: Agrawal, Garima, et autres
Publié: (2023)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
par: Zhu, Yanxu, et autres
Publié: (2024)
par: Zhu, Yanxu, et autres
Publié: (2024)
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
par: Gill, Waris, et autres
Publié: (2025)
par: Gill, Waris, et autres
Publié: (2025)
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
par: Hron, Jiri, et autres
Publié: (2024)
par: Hron, Jiri, et autres
Publié: (2024)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
par: Son, Guijin, et autres
Publié: (2023)
par: Son, Guijin, et autres
Publié: (2023)
Knowledge Verification to Nip Hallucination in the Bud
par: Wan, Fanqi, et autres
Publié: (2024)
par: Wan, Fanqi, et autres
Publié: (2024)
CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing
par: Hu, Yucheng, et autres
Publié: (2026)
par: Hu, Yucheng, et autres
Publié: (2026)
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
par: Yuan, Jiarui, et autres
Publié: (2026)
par: Yuan, Jiarui, et autres
Publié: (2026)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
par: Su, Zhaochen, et autres
Publié: (2024)
par: Su, Zhaochen, et autres
Publié: (2024)
nicolay-r at SemEval-2024 Task 3: Using Flan-T5 for Reasoning Emotion Cause in Conversations with Chain-of-Thought on Emotion States
par: Rusnachenko, Nicolay, et autres
Publié: (2024)
par: Rusnachenko, Nicolay, et autres
Publié: (2024)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
par: Hu, Nan, et autres
Publié: (2024)
par: Hu, Nan, et autres
Publié: (2024)
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
par: Kale, Sahil, et autres
Publié: (2025)
par: Kale, Sahil, et autres
Publié: (2025)
Documents similaires
-
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
par: Wang, Sizhe, et autres
Publié: (2024) -
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
par: Sansford, Hannah, et autres
Publié: (2024) -
D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs
par: Ding, Yue, et autres
Publié: (2025) -
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
par: Markowitz, Elan, et autres
Publié: (2025) -
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
par: Lavrinovics, Ernests, et autres
Publié: (2025)