RefChecker: Reference-based Fine-grained Hallucination Checker and Benchmark for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Xiangkun, Ru, Dongyu, Qiu, Lin, Guo, Qipeng, Zhang, Tianhang, Xu, Yang, Luo, Yun, Liu, Pengfei, Zhang, Yue, Zhang, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NCL-UoR at SemEval-2025 Task 3: Detecting Multilingual Hallucination and Related Observable Overgeneration Text Spans with Modified RefChecker and Modified SeflCheckGPT
von: Hong, Jiaying, et al.
Veröffentlicht: (2025)
von: Hong, Jiaying, et al.
Veröffentlicht: (2025)
Are Large Language Models Table-based Fact-Checkers?
von: Zhang, Hanwen, et al.
Veröffentlicht: (2024)
von: Zhang, Hanwen, et al.
Veröffentlicht: (2024)
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
von: Ru, Dongyu, et al.
Veröffentlicht: (2024)
von: Ru, Dongyu, et al.
Veröffentlicht: (2024)
Interactive DualChecker for Mitigating Hallucinations in Distilling Large Language Models
von: Wang, Meiyun, et al.
Veröffentlicht: (2024)
von: Wang, Meiyun, et al.
Veröffentlicht: (2024)
Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language Models
von: Min, Qingkai, et al.
Veröffentlicht: (2024)
von: Min, Qingkai, et al.
Veröffentlicht: (2024)
IoChecker
von: TNO
Veröffentlicht: (2025)
von: TNO
Veröffentlicht: (2025)
Can Language Models Learn to Skip Steps?
von: Liu, Tengxiao, et al.
Veröffentlicht: (2024)
von: Liu, Tengxiao, et al.
Veröffentlicht: (2024)
Write Your Own CodeChecker: An Automated Test-Driven Checker Development Approach with LLMs
von: Liu, Jun, et al.
Veröffentlicht: (2024)
von: Liu, Jun, et al.
Veröffentlicht: (2024)
Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker
von: Wu, Zongjian, et al.
Veröffentlicht: (2026)
von: Wu, Zongjian, et al.
Veröffentlicht: (2026)
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models
von: Li, Miaoran, et al.
Veröffentlicht: (2023)
von: Li, Miaoran, et al.
Veröffentlicht: (2023)
PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and Summarization
von: Luo, Yun, et al.
Veröffentlicht: (2024)
von: Luo, Yun, et al.
Veröffentlicht: (2024)
What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2026)
Billiards, Checkers, and Quadratic Reciprocity
von: Wästlund, Johan
Veröffentlicht: (2024)
von: Wästlund, Johan
Veröffentlicht: (2024)
The SCAN Statistical Model Checker
von: Ghiorzi, Enrico, et al.
Veröffentlicht: (2026)
von: Ghiorzi, Enrico, et al.
Veröffentlicht: (2026)
Oracle-Checker Scheme for Evaluating a Generative Large Language Model
von: Zeng, Yueling Jenny, et al.
Veröffentlicht: (2024)
von: Zeng, Yueling Jenny, et al.
Veröffentlicht: (2024)
LLMEffiChecker: Understanding and Testing Efficiency Degradation of Large Language Models
von: Feng, Xiaoning, et al.
Veröffentlicht: (2022)
von: Feng, Xiaoning, et al.
Veröffentlicht: (2022)
KNighter: Transforming Static Analysis with LLM-Synthesized Checkers
von: Yang, Chenyuan, et al.
Veröffentlicht: (2025)
von: Yang, Chenyuan, et al.
Veröffentlicht: (2025)
GenAI vs. Human Fact-Checkers: Accurate Ratings, Flawed Rationales
von: Tai, Yuehong Cassandra, et al.
Veröffentlicht: (2025)
von: Tai, Yuehong Cassandra, et al.
Veröffentlicht: (2025)
Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
A Lazy, Concurrent Convertibility Checker
von: Courant, Nathanaëlle, et al.
Veröffentlicht: (2025)
von: Courant, Nathanaëlle, et al.
Veröffentlicht: (2025)
Double Auctions: Formalization and Automated Checkers
von: Garg, Mohit, et al.
Veröffentlicht: (2024)
von: Garg, Mohit, et al.
Veröffentlicht: (2024)
Static Code Analysis with CodeChecker
von: Horvath, Gabor, et al.
Veröffentlicht: (2024)
von: Horvath, Gabor, et al.
Veröffentlicht: (2024)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
von: Sun, Yuhong, et al.
Veröffentlicht: (2024)
von: Sun, Yuhong, et al.
Veröffentlicht: (2024)
DocChecker: Bootstrapping Code Large Language Model for Detecting and Resolving Code-Comment Inconsistencies
von: Dau, Anh T. V., et al.
Veröffentlicht: (2023)
von: Dau, Anh T. V., et al.
Veröffentlicht: (2023)
Vibe Checker: Aligning Code Evaluation with Human Preference
von: Zhong, Ming, et al.
Veröffentlicht: (2025)
von: Zhong, Ming, et al.
Veröffentlicht: (2025)
MathMistake Checker: A Comprehensive Demonstration for Step-by-Step Math Problem Mistake Finding by Prompt-Guided LLMs
von: Zhang, Tianyang, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyang, et al.
Veröffentlicht: (2025)
GS-Checker: Tampering Localization for 3D Gaussian Splatting
von: Han, Haoliang, et al.
Veröffentlicht: (2025)
von: Han, Haoliang, et al.
Veröffentlicht: (2025)
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
von: Do, Quyet V., et al.
Veröffentlicht: (2024)
von: Do, Quyet V., et al.
Veröffentlicht: (2024)
foetus -- Termination Checker for Simple Functional Programs
von: Abel, Andreas
Veröffentlicht: (2024)
von: Abel, Andreas
Veröffentlicht: (2024)
AICCE: AI Driven Compliance Checker Engine
von: Rahman, Mohammad Wali Ur, et al.
Veröffentlicht: (2026)
von: Rahman, Mohammad Wali Ur, et al.
Veröffentlicht: (2026)
The rIC3 Hardware Model Checker
von: Su, Yuheng, et al.
Veröffentlicht: (2025)
von: Su, Yuheng, et al.
Veröffentlicht: (2025)
A Model Checker for Natural Strategic Ability
von: Aruta, Marco, et al.
Veröffentlicht: (2024)
von: Aruta, Marco, et al.
Veröffentlicht: (2024)
Variants of Conway Checkers and k-nacci Jumping
von: Bruda, Glenn, et al.
Veröffentlicht: (2024)
von: Bruda, Glenn, et al.
Veröffentlicht: (2024)
Renal Denervation—“Gizmo Idolatry” Fact Checker
von: Markus P. Schlaich, et al.
Veröffentlicht: (2025)
von: Markus P. Schlaich, et al.
Veröffentlicht: (2025)
AutoTAMP: Autoregressive Task and Motion Planning with LLMs as Translators and Checkers
von: Chen, Yongchao, et al.
Veröffentlicht: (2023)
von: Chen, Yongchao, et al.
Veröffentlicht: (2023)
State Space Estimation for DPOR-based Model Checkers(Extended Version)
von: Balasubramanian, A. R., et al.
Veröffentlicht: (2025)
von: Balasubramanian, A. R., et al.
Veröffentlicht: (2025)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
von: Sakai, Yusuke, et al.
Veröffentlicht: (2026)
von: Sakai, Yusuke, et al.
Veröffentlicht: (2026)
Characterizing AI Fact-Checkers and Their Contributions on Community Notes
von: Gong, Yilin, et al.
Veröffentlicht: (2026)
von: Gong, Yilin, et al.
Veröffentlicht: (2026)
ExeChecker: Where Did I Go Wrong?
von: Gu, Yiwen, et al.
Veröffentlicht: (2024)
von: Gu, Yiwen, et al.
Veröffentlicht: (2024)
Can Community Notes Replace Professional Fact-Checkers?
von: Borenstein, Nadav, et al.
Veröffentlicht: (2025)
von: Borenstein, Nadav, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NCL-UoR at SemEval-2025 Task 3: Detecting Multilingual Hallucination and Related Observable Overgeneration Text Spans with Modified RefChecker and Modified SeflCheckGPT
von: Hong, Jiaying, et al.
Veröffentlicht: (2025) -
Are Large Language Models Table-based Fact-Checkers?
von: Zhang, Hanwen, et al.
Veröffentlicht: (2024) -
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
von: Ru, Dongyu, et al.
Veröffentlicht: (2024) -
Interactive DualChecker for Mitigating Hallucinations in Distilling Large Language Models
von: Wang, Meiyun, et al.
Veröffentlicht: (2024) -
Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language Models
von: Min, Qingkai, et al.
Veröffentlicht: (2024)