Saved in:
Bibliographic Details
Main Authors: Strauss, Ilan, Yang, Jangho, O'Reilly, Tim, Rosenblat, Sruly, Moure, Isobel
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.00838
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912515611426816
author Strauss, Ilan
Yang, Jangho
O'Reilly, Tim
Rosenblat, Sruly
Moure, Isobel
author_facet Strauss, Ilan
Yang, Jangho
O'Reilly, Tim
Rosenblat, Sruly
Moure, Isobel
contents Web-enabled LLMs frequently answer queries without crediting the web pages they consume, creating an "attribution gap" - the difference between relevant URLs read and those actually cited. Drawing on approximately 14,000 real-world LMArena conversation logs with search-enabled LLM systems, we document three exploitation patterns: 1) No Search: 34% of Google Gemini and 24% of OpenAI GPT-4o responses are generated without explicitly fetching any online content; 2) No citation: Gemini provides no clickable citation source in 92% of answers; 3) High-volume, low-credit: Perplexity's Sonar visits approximately 10 relevant pages per query but cites only three to four. A negative binomial hurdle model shows that the average query answered by Gemini or Sonar leaves about 3 relevant websites uncited, whereas GPT-4o's tiny uncited gap is best explained by its selective log disclosures rather than by better attribution. Citation efficiency - extra citations provided per additional relevant web page visited - varies widely across models, from 0.19 to 0.45 on identical queries, underscoring that retrieval design, not technical limits, shapes ecosystem impact. We recommend a transparent LLM search architecture based on standardized telemetry and full disclosure of search traces and citation logs.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00838
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Attribution Crisis in LLM Search Results
Strauss, Ilan
Yang, Jangho
O'Reilly, Tim
Rosenblat, Sruly
Moure, Isobel
Digital Libraries
Artificial Intelligence
Computation and Language
Web-enabled LLMs frequently answer queries without crediting the web pages they consume, creating an "attribution gap" - the difference between relevant URLs read and those actually cited. Drawing on approximately 14,000 real-world LMArena conversation logs with search-enabled LLM systems, we document three exploitation patterns: 1) No Search: 34% of Google Gemini and 24% of OpenAI GPT-4o responses are generated without explicitly fetching any online content; 2) No citation: Gemini provides no clickable citation source in 92% of answers; 3) High-volume, low-credit: Perplexity's Sonar visits approximately 10 relevant pages per query but cites only three to four. A negative binomial hurdle model shows that the average query answered by Gemini or Sonar leaves about 3 relevant websites uncited, whereas GPT-4o's tiny uncited gap is best explained by its selective log disclosures rather than by better attribution. Citation efficiency - extra citations provided per additional relevant web page visited - varies widely across models, from 0.19 to 0.45 on identical queries, underscoring that retrieval design, not technical limits, shapes ecosystem impact. We recommend a transparent LLM search architecture based on standardized telemetry and full disclosure of search traces and citation logs.
title The Attribution Crisis in LLM Search Results
topic Digital Libraries
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.00838