SelfCheck-Eval: A Multi-Module Framework for Zero-Resource Hallucination Detection in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Muhammed, Diyana, Tuccari, Giusy Giulia, Rabby, Gollam, Auer, Sören, Vahdati, Sahar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees
by: Rabby, Gollam, et al.
Published: (2025)
by: Rabby, Gollam, et al.
Published: (2025)
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
by: Keya, Farhana, et al.
Published: (2025)
by: Keya, Farhana, et al.
Published: (2025)
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
by: Gaddipati, Sasi Kiran, et al.
Published: (2026)
by: Gaddipati, Sasi Kiran, et al.
Published: (2026)
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
by: Rabby, Gollam, et al.
Published: (2024)
by: Rabby, Gollam, et al.
Published: (2024)
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
by: Rabby, Gollam, et al.
Published: (2024)
by: Rabby, Gollam, et al.
Published: (2024)
NeuroSym-BioCAT: Leveraging Neuro-Symbolic Methods for Biomedical Scholarly Document Categorization and Question Answering
by: Zamil, Parvez, et al.
Published: (2024)
by: Zamil, Parvez, et al.
Published: (2024)
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
by: Gaddipati, Sasi Kiran, et al.
Published: (2025)
by: Gaddipati, Sasi Kiran, et al.
Published: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination
by: Li, Zhuo, et al.
Published: (2026)
by: Li, Zhuo, et al.
Published: (2026)
Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
by: Liu, Qiang, et al.
Published: (2025)
by: Liu, Qiang, et al.
Published: (2025)
Computational Fact-Checking of Online Discourse: Scoring scientific accuracy in climate change related news articles
by: Wittenborg, Tim, et al.
Published: (2025)
by: Wittenborg, Tim, et al.
Published: (2025)
Large Language Models as Evaluators for Scientific Synthesis
by: Evans, Julia, et al.
Published: (2024)
by: Evans, Julia, et al.
Published: (2024)
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
by: Li, Miaoran, et al.
Published: (2023)
by: Li, Miaoran, et al.
Published: (2023)
Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
by: Sinha, Samridhi Raj, et al.
Published: (2025)
by: Sinha, Samridhi Raj, et al.
Published: (2025)
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
by: Ahmed, Tawsif, et al.
Published: (2025)
by: Ahmed, Tawsif, et al.
Published: (2025)
Zero-Shot Multi-task Hallucination Detection
by: Bhamidipati, Patanjali, et al.
Published: (2024)
by: Bhamidipati, Patanjali, et al.
Published: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)
by: Hosseini, Mohammad, et al.
Published: (2025)
Framework for Hallucination Detection in Large Language Models
by: Dheeraj Sundaragiri, et al.
Published: (2026)
by: Dheeraj Sundaragiri, et al.
Published: (2026)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
Large Language Models for Scientific Information Extraction: An Empirical Study for Virology
by: Shamsabadi, Mahsa, et al.
Published: (2024)
by: Shamsabadi, Mahsa, et al.
Published: (2024)
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
by: Kummer, Cornelius, et al.
Published: (2026)
by: Kummer, Cornelius, et al.
Published: (2026)
LLMs4Synthesis: Leveraging Large Language Models for Scientific Synthesis
by: Giglou, Hamed Babaei, et al.
Published: (2024)
by: Giglou, Hamed Babaei, et al.
Published: (2024)
LLMs4OL 2024 Overview: The 1st Large Language Models for Ontology Learning Challenge
by: Giglou, Hamed Babaei, et al.
Published: (2024)
by: Giglou, Hamed Babaei, et al.
Published: (2024)
Sanity Checks for Long-Form Hallucination Detection
by: Zollicoffer, Geigh, et al.
Published: (2026)
by: Zollicoffer, Geigh, et al.
Published: (2026)
OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models
by: Sun, Chongren, et al.
Published: (2025)
by: Sun, Chongren, et al.
Published: (2025)
CiteFusion: An Ensemble Framework for Citation Intent Classification Harnessing Dual-Model Binary Couples and SHAP Analyses
by: Paolini, Lorenzo, et al.
Published: (2024)
by: Paolini, Lorenzo, et al.
Published: (2024)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
by: Lee, DongGeon, et al.
Published: (2025)
by: Lee, DongGeon, et al.
Published: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
by: Sawczyn, Albert, et al.
Published: (2025)
by: Sawczyn, Albert, et al.
Published: (2025)
InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated Answers
by: Yehuda, Yakir, et al.
Published: (2024)
by: Yehuda, Yakir, et al.
Published: (2024)
AILS-NTUA at SemEval-2025 Task 3: Leveraging Large Language Models and Translation Strategies for Multilingual Hallucination Detection
by: Karkani, Dimitra, et al.
Published: (2025)
by: Karkani, Dimitra, et al.
Published: (2025)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
by: Zhu, Zhiying, et al.
Published: (2024)
by: Zhu, Zhiying, et al.
Published: (2024)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
by: Zhang, Kaichen, et al.
Published: (2024)
by: Zhang, Kaichen, et al.
Published: (2024)
SciCom Wiki: Fact-Checking and FAIR Knowledge Distribution for Scientific Videos and Podcasts
by: Wittenborg, Tim, et al.
Published: (2025)
by: Wittenborg, Tim, et al.
Published: (2025)
Aligning Knowledge Graphs and Language Models for Factual Accuracy
by: Nishat, Nur A Zarin, et al.
Published: (2025)
by: Nishat, Nur A Zarin, et al.
Published: (2025)
SHROOM-INDElab at SemEval-2024 Task 6: Zero- and Few-Shot LLM-Based Classification for Hallucination Detection
by: Allen, Bradley P., et al.
Published: (2024)
by: Allen, Bradley P., et al.
Published: (2024)
A Multiple-Fill-in-the-Blank Exam Approach for Enhancing Zero-Resource Hallucination Detection in Large Language Models
by: Munakata, Satoshi, et al.
Published: (2024)
by: Munakata, Satoshi, et al.
Published: (2024)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
by: Mündler, Niels, et al.
Published: (2023)
by: Mündler, Niels, et al.
Published: (2023)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
by: Zhang, Jiawei, et al.
Published: (2024)
by: Zhang, Jiawei, et al.
Published: (2024)
Similar Items
-
Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees
by: Rabby, Gollam, et al.
Published: (2025) -
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
by: Keya, Farhana, et al.
Published: (2025) -
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
by: Gaddipati, Sasi Kiran, et al.
Published: (2026) -
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
by: Rabby, Gollam, et al.
Published: (2024) -
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
by: Rabby, Gollam, et al.
Published: (2024)