LLM hallucinations in the wild: Large-scale evidence from non-existent citations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Zhenyue, Wang, Yihe, Stuart, Toby, De Vaan, Mathijs, Ginsparg, Paul, Yin, Yian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914544733913088
author Zhao, Zhenyue
Wang, Yihe
Stuart, Toby
De Vaan, Mathijs
Ginsparg, Paul
Yin, Yian
author_facet Zhao, Zhenyue
Wang, Yihe
Stuart, Toby
De Vaan, Mathijs
Ginsparg, Paul
Yin, Yian
contents Large language models (LLMs) are known to generate plausible but false information across a wide range of contexts, yet the real-world magnitude and consequences of this hallucination problem remain poorly understood. Here we leverage a uniquely verifiable object - scientific citations - to audit 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central. We find a sharp rise in non-existent references following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone. These errors are diffusely embedded across many papers but especially pronounced in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI-assisted writing, and among small and early-career author teams. At the same time, hallucinated references disproportionately assign credit to already prominent and male scholars, suggesting that LLM-generated errors may reinforce existing inequities in scientific recognition. Preprint moderation and journal publication processes capture only a fraction of these errors, suggesting that the spread of hallucinated content has outpaced existing safeguards. Together, these findings demonstrate that LLM hallucinations are infiltrating knowledge production at scale, threatening both the reliability and equity of future scientific discovery as human and AI systems draw on the existing literature.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07723
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LLM hallucinations in the wild: Large-scale evidence from non-existent citations
Zhao, Zhenyue
Wang, Yihe
Stuart, Toby
De Vaan, Mathijs
Ginsparg, Paul
Yin, Yian
Digital Libraries
Artificial Intelligence
Computers and Society
Physics and Society
Large language models (LLMs) are known to generate plausible but false information across a wide range of contexts, yet the real-world magnitude and consequences of this hallucination problem remain poorly understood. Here we leverage a uniquely verifiable object - scientific citations - to audit 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central. We find a sharp rise in non-existent references following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone. These errors are diffusely embedded across many papers but especially pronounced in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI-assisted writing, and among small and early-career author teams. At the same time, hallucinated references disproportionately assign credit to already prominent and male scholars, suggesting that LLM-generated errors may reinforce existing inequities in scientific recognition. Preprint moderation and journal publication processes capture only a fraction of these errors, suggesting that the spread of hallucinated content has outpaced existing safeguards. Together, these findings demonstrate that LLM hallucinations are infiltrating knowledge production at scale, threatening both the reliability and equity of future scientific discovery as human and AI systems draw on the existing literature.
title LLM hallucinations in the wild: Large-scale evidence from non-existent citations
topic Digital Libraries
Artificial Intelligence
Computers and Society
Physics and Society
url https://arxiv.org/abs/2605.07723