SelfCheck-Eval: A Multi-Module Framework for Zero-Resource Hallucination Detection in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Muhammed, Diyana, Tuccari, Giusy Giulia, Rabby, Gollam, Auer, Sören, Vahdati, Sahar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees
von: Rabby, Gollam, et al.
Veröffentlicht: (2025)
von: Rabby, Gollam, et al.
Veröffentlicht: (2025)
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
von: Keya, Farhana, et al.
Veröffentlicht: (2025)
von: Keya, Farhana, et al.
Veröffentlicht: (2025)
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
von: Gaddipati, Sasi Kiran, et al.
Veröffentlicht: (2026)
von: Gaddipati, Sasi Kiran, et al.
Veröffentlicht: (2026)
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
von: Rabby, Gollam, et al.
Veröffentlicht: (2024)
von: Rabby, Gollam, et al.
Veröffentlicht: (2024)
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
von: Rabby, Gollam, et al.
Veröffentlicht: (2024)
von: Rabby, Gollam, et al.
Veröffentlicht: (2024)
NeuroSym-BioCAT: Leveraging Neuro-Symbolic Methods for Biomedical Scholarly Document Categorization and Question Answering
von: Zamil, Parvez, et al.
Veröffentlicht: (2024)
von: Zamil, Parvez, et al.
Veröffentlicht: (2024)
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
von: Schuhmann, Christoph, et al.
Veröffentlicht: (2025)
von: Schuhmann, Christoph, et al.
Veröffentlicht: (2025)
AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
von: Gaddipati, Sasi Kiran, et al.
Veröffentlicht: (2025)
von: Gaddipati, Sasi Kiran, et al.
Veröffentlicht: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination
von: Li, Zhuo, et al.
Veröffentlicht: (2026)
von: Li, Zhuo, et al.
Veröffentlicht: (2026)
Project Alexandria: Towards Freeing Scientific Knowledge from Copyright Burdens via LLMs
von: Schuhmann, Christoph, et al.
Veröffentlicht: (2025)
von: Schuhmann, Christoph, et al.
Veröffentlicht: (2025)
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
Computational Fact-Checking of Online Discourse: Scoring scientific accuracy in climate change related news articles
von: Wittenborg, Tim, et al.
Veröffentlicht: (2025)
von: Wittenborg, Tim, et al.
Veröffentlicht: (2025)
Large Language Models as Evaluators for Scientific Synthesis
von: Evans, Julia, et al.
Veröffentlicht: (2024)
von: Evans, Julia, et al.
Veröffentlicht: (2024)
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
von: Li, Miaoran, et al.
Veröffentlicht: (2023)
von: Li, Miaoran, et al.
Veröffentlicht: (2023)
Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
von: Sinha, Samridhi Raj, et al.
Veröffentlicht: (2025)
von: Sinha, Samridhi Raj, et al.
Veröffentlicht: (2025)
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
von: Ahmed, Tawsif, et al.
Veröffentlicht: (2025)
von: Ahmed, Tawsif, et al.
Veröffentlicht: (2025)
Zero-Shot Multi-task Hallucination Detection
von: Bhamidipati, Patanjali, et al.
Veröffentlicht: (2024)
von: Bhamidipati, Patanjali, et al.
Veröffentlicht: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
von: Hosseini, Mohammad, et al.
Veröffentlicht: (2025)
Framework for Hallucination Detection in Large Language Models
von: Dheeraj Sundaragiri, et al.
Veröffentlicht: (2026)
von: Dheeraj Sundaragiri, et al.
Veröffentlicht: (2026)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
Large Language Models for Scientific Information Extraction: An Empirical Study for Virology
von: Shamsabadi, Mahsa, et al.
Veröffentlicht: (2024)
von: Shamsabadi, Mahsa, et al.
Veröffentlicht: (2024)
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
von: Kummer, Cornelius, et al.
Veröffentlicht: (2026)
von: Kummer, Cornelius, et al.
Veröffentlicht: (2026)
LLMs4Synthesis: Leveraging Large Language Models for Scientific Synthesis
von: Giglou, Hamed Babaei, et al.
Veröffentlicht: (2024)
von: Giglou, Hamed Babaei, et al.
Veröffentlicht: (2024)
LLMs4OL 2024 Overview: The 1st Large Language Models for Ontology Learning Challenge
von: Giglou, Hamed Babaei, et al.
Veröffentlicht: (2024)
von: Giglou, Hamed Babaei, et al.
Veröffentlicht: (2024)
Sanity Checks for Long-Form Hallucination Detection
von: Zollicoffer, Geigh, et al.
Veröffentlicht: (2026)
von: Zollicoffer, Geigh, et al.
Veröffentlicht: (2026)
OnionEval: An Unified Evaluation of Fact-conflicting Hallucination for Small-Large Language Models
von: Sun, Chongren, et al.
Veröffentlicht: (2025)
von: Sun, Chongren, et al.
Veröffentlicht: (2025)
CiteFusion: An Ensemble Framework for Citation Intent Classification Harnessing Dual-Model Binary Couples and SHAP Analyses
von: Paolini, Lorenzo, et al.
Veröffentlicht: (2024)
von: Paolini, Lorenzo, et al.
Veröffentlicht: (2024)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
von: Lee, DongGeon, et al.
Veröffentlicht: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated Answers
von: Yehuda, Yakir, et al.
Veröffentlicht: (2024)
von: Yehuda, Yakir, et al.
Veröffentlicht: (2024)
AILS-NTUA at SemEval-2025 Task 3: Leveraging Large Language Models and Translation Strategies for Multilingual Hallucination Detection
von: Karkani, Dimitra, et al.
Veröffentlicht: (2025)
von: Karkani, Dimitra, et al.
Veröffentlicht: (2025)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
von: Zhu, Zhiying, et al.
Veröffentlicht: (2024)
von: Zhu, Zhiying, et al.
Veröffentlicht: (2024)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
SciCom Wiki: Fact-Checking and FAIR Knowledge Distribution for Scientific Videos and Podcasts
von: Wittenborg, Tim, et al.
Veröffentlicht: (2025)
von: Wittenborg, Tim, et al.
Veröffentlicht: (2025)
Aligning Knowledge Graphs and Language Models for Factual Accuracy
von: Nishat, Nur A Zarin, et al.
Veröffentlicht: (2025)
von: Nishat, Nur A Zarin, et al.
Veröffentlicht: (2025)
SHROOM-INDElab at SemEval-2024 Task 6: Zero- and Few-Shot LLM-Based Classification for Hallucination Detection
von: Allen, Bradley P., et al.
Veröffentlicht: (2024)
von: Allen, Bradley P., et al.
Veröffentlicht: (2024)
A Multiple-Fill-in-the-Blank Exam Approach for Enhancing Zero-Resource Hallucination Detection in Large Language Models
von: Munakata, Satoshi, et al.
Veröffentlicht: (2024)
von: Munakata, Satoshi, et al.
Veröffentlicht: (2024)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees
von: Rabby, Gollam, et al.
Veröffentlicht: (2025) -
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings
von: Keya, Farhana, et al.
Veröffentlicht: (2025) -
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
von: Gaddipati, Sasi Kiran, et al.
Veröffentlicht: (2026) -
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
von: Rabby, Gollam, et al.
Veröffentlicht: (2024) -
Fine-tuning and Prompt Engineering with Cognitive Knowledge Graphs for Scholarly Knowledge Organization
von: Rabby, Gollam, et al.
Veröffentlicht: (2024)