Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: DeVerna, Matthew R., Yang, Kai-Cheng, Yan, Harry Yaojun, Menczer, Filippo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918216276639744
author DeVerna, Matthew R.
Yang, Kai-Cheng
Yan, Harry Yaojun
Menczer, Filippo
author_facet DeVerna, Matthew R.
Yang, Kai-Cheng
Yan, Harry Yaojun
Menczer, Filippo
contents Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results. As mainstream chatbots increasingly ship with reasoning capabilities and web search tools -- and millions of users already rely on them for verification -- rigorous evaluation is urgent. We evaluate 15 recent LLMs from OpenAI, Google, Meta, and DeepSeek on more than 6,000 claims fact-checked by PolitiFact, comparing standard models with reasoning- and web-search variants. Standard models perform poorly, reasoning offers minimal benefits, and web search provides only moderate gains, despite fact-checks being available on the web. In contrast, a curated RAG system using PolitiFact summaries improved macro F1 by 233% on average across model variants. These findings suggest that giving models access to curated high-quality context is a promising path for automated fact-checking.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
DeVerna, Matthew R.
Yang, Kai-Cheng
Yan, Harry Yaojun
Menczer, Filippo
Computation and Language
Computers and Society
Information Retrieval
Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results. As mainstream chatbots increasingly ship with reasoning capabilities and web search tools -- and millions of users already rely on them for verification -- rigorous evaluation is urgent. We evaluate 15 recent LLMs from OpenAI, Google, Meta, and DeepSeek on more than 6,000 claims fact-checked by PolitiFact, comparing standard models with reasoning- and web-search variants. Standard models perform poorly, reasoning offers minimal benefits, and web search provides only moderate gains, despite fact-checks being available on the web. In contrast, a curated RAG system using PolitiFact summaries improved macro F1 by 233% on average across model variants. These findings suggest that giving models access to curated high-quality context is a promising path for automated fact-checking.
title Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
topic Computation and Language
Computers and Society
Information Retrieval
url https://arxiv.org/abs/2511.18749