Scaling Truth: The Confidence Paradox in AI Fact-Checking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qazi, Ihsan A., Khan, Zohaib, Ghani, Abdullah, Raza, Agha A., Qazi, Zafar A., Sajjad, Wassay, Ali, Ayesha, Javaid, Asher, Sohail, Muhammad Abdullah, Azeemi, Abdul H.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914031504195584
author Qazi, Ihsan A.
Khan, Zohaib
Ghani, Abdullah
Raza, Agha A.
Qazi, Zafar A.
Sajjad, Wassay
Ali, Ayesha
Javaid, Asher
Sohail, Muhammad Abdullah
Azeemi, Abdul H.
author_facet Qazi, Ihsan A.
Khan, Zohaib
Ghani, Abdullah
Raza, Agha A.
Qazi, Zafar A.
Sajjad, Wassay
Ali, Ayesha
Javaid, Asher
Sohail, Muhammad Abdullah
Azeemi, Abdul H.
contents The rise of misinformation underscores the need for scalable and reliable fact-checking solutions. Large language models (LLMs) hold promise in automating fact verification, yet their effectiveness across global contexts remains uncertain. We systematically evaluate nine established LLMs across multiple categories (open/closed-source, multiple sizes, diverse architectures, reasoning-based) using 5,000 claims previously assessed by 174 professional fact-checking organizations across 47 languages. Our methodology tests model generalizability on claims postdating training cutoffs and four prompting strategies mirroring both citizen and professional fact-checker interactions, with over 240,000 human annotations as ground truth. Findings reveal a concerning pattern resembling the Dunning-Kruger effect: smaller, accessible models show high confidence despite lower accuracy, while larger models demonstrate higher accuracy but lower confidence. This risks systemic bias in information verification, as resource-constrained organizations typically use smaller models. Performance gaps are most pronounced for non-English languages and claims originating from the Global South, threatening to widen existing information inequalities. These results establish a multilingual benchmark for future research and provide an evidence base for policy aimed at ensuring equitable access to trustworthy, AI-assisted fact-checking.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Truth: The Confidence Paradox in AI Fact-Checking
Qazi, Ihsan A.
Khan, Zohaib
Ghani, Abdullah
Raza, Agha A.
Qazi, Zafar A.
Sajjad, Wassay
Ali, Ayesha
Javaid, Asher
Sohail, Muhammad Abdullah
Azeemi, Abdul H.
Social and Information Networks
Artificial Intelligence
Computation and Language
Computers and Society
The rise of misinformation underscores the need for scalable and reliable fact-checking solutions. Large language models (LLMs) hold promise in automating fact verification, yet their effectiveness across global contexts remains uncertain. We systematically evaluate nine established LLMs across multiple categories (open/closed-source, multiple sizes, diverse architectures, reasoning-based) using 5,000 claims previously assessed by 174 professional fact-checking organizations across 47 languages. Our methodology tests model generalizability on claims postdating training cutoffs and four prompting strategies mirroring both citizen and professional fact-checker interactions, with over 240,000 human annotations as ground truth. Findings reveal a concerning pattern resembling the Dunning-Kruger effect: smaller, accessible models show high confidence despite lower accuracy, while larger models demonstrate higher accuracy but lower confidence. This risks systemic bias in information verification, as resource-constrained organizations typically use smaller models. Performance gaps are most pronounced for non-English languages and claims originating from the Global South, threatening to widen existing information inequalities. These results establish a multilingual benchmark for future research and provide an evidence base for policy aimed at ensuring equitable access to trustworthy, AI-assisted fact-checking.
title Scaling Truth: The Confidence Paradox in AI Fact-Checking
topic Social and Information Networks
Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2509.08803