KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
Fuente:
arXiv
Salvato in:
| Autori principali: | Robertson, Alex, Liang, Huizhi, Gani, Mahbub, Kumar, Rohit, Rajamohan, Srijith |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
di: Wang, Sizhe, et al.
Pubblicazione: (2024)
di: Wang, Sizhe, et al.
Pubblicazione: (2024)
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
di: Sansford, Hannah, et al.
Pubblicazione: (2024)
D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs
di: Ding, Yue, et al.
Pubblicazione: (2025)
di: Ding, Yue, et al.
Pubblicazione: (2025)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
di: Lavrinovics, Ernests, et al.
Pubblicazione: (2025)
di: Lavrinovics, Ernests, et al.
Pubblicazione: (2025)
Training for Compositional Sensitivity Reduces Dense Retrieval Generalization
di: Ralev, Radoslav, et al.
Pubblicazione: (2026)
di: Ralev, Radoslav, et al.
Pubblicazione: (2026)
MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models
di: Liu, Weixin, et al.
Pubblicazione: (2026)
di: Liu, Weixin, et al.
Pubblicazione: (2026)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
di: Kumar, Mahesh, et al.
Pubblicazione: (2026)
di: Kumar, Mahesh, et al.
Pubblicazione: (2026)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
From Hallucinations to Facts: Enhancing Language Models with Curated Knowledge Graphs
di: Joshi, Ratnesh Kumar, et al.
Pubblicazione: (2024)
di: Joshi, Ratnesh Kumar, et al.
Pubblicazione: (2024)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
di: Ghanem, Hussam, et al.
Pubblicazione: (2025)
di: Ghanem, Hussam, et al.
Pubblicazione: (2025)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
di: Chen, Yongrui, et al.
Pubblicazione: (2025)
di: Chen, Yongrui, et al.
Pubblicazione: (2025)
CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders
di: Gulko, Alex, et al.
Pubblicazione: (2025)
di: Gulko, Alex, et al.
Pubblicazione: (2025)
Combining LLMs and Knowledge Graphs to Reduce Hallucinations in Question Answering
di: Pusch, Larissa, et al.
Pubblicazione: (2024)
di: Pusch, Larissa, et al.
Pubblicazione: (2024)
LastingBench: Defend Benchmarks Against Knowledge Leakage
di: Fang, Yixiong, et al.
Pubblicazione: (2025)
di: Fang, Yixiong, et al.
Pubblicazione: (2025)
Characterizing Knowledge Graph Tasks in LLM Benchmarks Using Cognitive Complexity Frameworks
di: Todorovikj, Sara, et al.
Pubblicazione: (2025)
di: Todorovikj, Sara, et al.
Pubblicazione: (2025)
Agent-as-a-Graph: Knowledge Graph-Based Tool and Agent Retrieval for LLM Multi-Agent Systems
di: Nizar, Faheem, et al.
Pubblicazione: (2025)
di: Nizar, Faheem, et al.
Pubblicazione: (2025)
The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
di: Zhang, Yuji, et al.
Pubblicazione: (2025)
LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoning
di: Xing, Wang, et al.
Pubblicazione: (2026)
di: Xing, Wang, et al.
Pubblicazione: (2026)
Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
di: Zaporojets, Klim, et al.
Pubblicazione: (2025)
di: Zaporojets, Klim, et al.
Pubblicazione: (2025)
MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge
di: Fu, Baochen, et al.
Pubblicazione: (2026)
di: Fu, Baochen, et al.
Pubblicazione: (2026)
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
di: Du, Yuntao, et al.
Pubblicazione: (2025)
di: Du, Yuntao, et al.
Pubblicazione: (2025)
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
di: Xie, Pengzhen, et al.
Pubblicazione: (2025)
di: Xie, Pengzhen, et al.
Pubblicazione: (2025)
On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
di: Kalavasis, Alkis, et al.
Pubblicazione: (2024)
di: Kalavasis, Alkis, et al.
Pubblicazione: (2024)
Grounding LLM Reasoning with Knowledge Graphs
di: Amayuelas, Alfonso, et al.
Pubblicazione: (2025)
di: Amayuelas, Alfonso, et al.
Pubblicazione: (2025)
Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
di: Agrawal, Garima, et al.
Pubblicazione: (2023)
di: Agrawal, Garima, et al.
Pubblicazione: (2023)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
di: Gill, Waris, et al.
Pubblicazione: (2025)
di: Gill, Waris, et al.
Pubblicazione: (2025)
Training Language Models on the Knowledge Graph: Insights on Hallucinations and Their Detectability
di: Hron, Jiri, et al.
Pubblicazione: (2024)
di: Hron, Jiri, et al.
Pubblicazione: (2024)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
di: Son, Guijin, et al.
Pubblicazione: (2023)
di: Son, Guijin, et al.
Pubblicazione: (2023)
Knowledge Verification to Nip Hallucination in the Bud
di: Wan, Fanqi, et al.
Pubblicazione: (2024)
di: Wan, Fanqi, et al.
Pubblicazione: (2024)
CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing
di: Hu, Yucheng, et al.
Pubblicazione: (2026)
di: Hu, Yucheng, et al.
Pubblicazione: (2026)
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
di: Yuan, Jiarui, et al.
Pubblicazione: (2026)
di: Yuan, Jiarui, et al.
Pubblicazione: (2026)
ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM
di: Su, Zhaochen, et al.
Pubblicazione: (2024)
di: Su, Zhaochen, et al.
Pubblicazione: (2024)
nicolay-r at SemEval-2024 Task 3: Using Flan-T5 for Reasoning Emotion Cause in Conversations with Chain-of-Thought on Emotion States
di: Rusnachenko, Nicolay, et al.
Pubblicazione: (2024)
di: Rusnachenko, Nicolay, et al.
Pubblicazione: (2024)
Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs
di: Hu, Nan, et al.
Pubblicazione: (2024)
di: Hu, Nan, et al.
Pubblicazione: (2024)
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
di: Kale, Sahil, et al.
Pubblicazione: (2025)
di: Kale, Sahil, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
di: Wang, Sizhe, et al.
Pubblicazione: (2024) -
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework
di: Sansford, Hannah, et al.
Pubblicazione: (2024) -
D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs
di: Ding, Yue, et al.
Pubblicazione: (2025) -
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
di: Markowitz, Elan, et al.
Pubblicazione: (2025) -
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
di: Lavrinovics, Ernests, et al.
Pubblicazione: (2025)