RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Shuo, Dai, Yuqin, Wang, Guoqing, Zheng, Xinran, Xu, Jinfeng, Li, Jinze, Ying, Zhenzhe, Wang, Weiqiang, Ngai, Edith C. H.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911006544887808
author Yang, Shuo
Dai, Yuqin
Wang, Guoqing
Zheng, Xinran
Xu, Jinfeng
Li, Jinze
Ying, Zhenzhe
Wang, Weiqiang
Ngai, Edith C. H.
author_facet Yang, Shuo
Dai, Yuqin
Wang, Guoqing
Zheng, Xinran
Xu, Jinfeng
Li, Jinze
Ying, Zhenzhe
Wang, Weiqiang
Ngai, Edith C. H.
contents Large Language Models (LLMs) hold significant potential for advancing fact-checking by leveraging their capabilities in reasoning, evidence retrieval, and explanation generation. However, existing benchmarks fail to comprehensively evaluate LLMs and Multimodal Large Language Models (MLLMs) in realistic misinformation scenarios. To bridge this gap, we introduce RealFactBench, a comprehensive benchmark designed to assess the fact-checking capabilities of LLMs and MLLMs across diverse real-world tasks, including Knowledge Validation, Rumor Detection, and Event Verification. RealFactBench consists of 6K high-quality claims drawn from authoritative sources, encompassing multimodal content and diverse domains. Our evaluation framework further introduces the Unknown Rate (UnR) metric, enabling a more nuanced assessment of models' ability to handle uncertainty and balance between over-conservatism and over-confidence. Extensive experiments on 7 representative LLMs and 4 MLLMs reveal their limitations in real-world fact-checking and offer valuable insights for further research. RealFactBench is publicly available at https://github.com/kalendsyang/RealFactBench.git.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
Yang, Shuo
Dai, Yuqin
Wang, Guoqing
Zheng, Xinran
Xu, Jinfeng
Li, Jinze
Ying, Zhenzhe
Wang, Weiqiang
Ngai, Edith C. H.
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) hold significant potential for advancing fact-checking by leveraging their capabilities in reasoning, evidence retrieval, and explanation generation. However, existing benchmarks fail to comprehensively evaluate LLMs and Multimodal Large Language Models (MLLMs) in realistic misinformation scenarios. To bridge this gap, we introduce RealFactBench, a comprehensive benchmark designed to assess the fact-checking capabilities of LLMs and MLLMs across diverse real-world tasks, including Knowledge Validation, Rumor Detection, and Event Verification. RealFactBench consists of 6K high-quality claims drawn from authoritative sources, encompassing multimodal content and diverse domains. Our evaluation framework further introduces the Unknown Rate (UnR) metric, enabling a more nuanced assessment of models' ability to handle uncertainty and balance between over-conservatism and over-confidence. Extensive experiments on 7 representative LLMs and 4 MLLMs reveal their limitations in real-world fact-checking and offer valuable insights for further research. RealFactBench is publicly available at https://github.com/kalendsyang/RealFactBench.git.
title RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.12538